Conference Agenda
| Session | ||
09 SES 03 A: Fairness and Validity in High-Stakes Testing (Part 2): Validity in Practice: How Standards, Interpretations, and AI-Enabled Operations Shape Fair High-Stakes Testing
Symposium | ||
| Presentations | ||
09. Assessment, Evaluation, Testing and Measurement
Symposium Fairness and Validity in High-Stakes Testing (Part 2): "Validity in Practice: How Standards, Interpretations, and AI-Enabled Operations Shape Fair High-Stakes Testing The overarching theme is centred around ensuring fairness, validity, and reliability (Kane 2013, Messick, 1989) in high-stakes educational assessments, with a strong emphasis on how different factors (such as item format, gender, language background, socio-economic status, scoring methods, and standard-setting practices) can introduce bias or inequity. High-stakes test are considered “assessments that have serious consequences” for participants (Stobart & Eggen, 2012, p.1) . They typically involve national assessments or external examinations used for admissions to a higher level of education or that otherwise have a significant impact on students’ educational careers and other life choices. With this in mind, fairness and validity are corner concepts of any modern high-stakes testing. Four general aspects of fairness are the equivalent opportunity of test-takers in the testing process, the lack of measurement bias, the access to the construct(s) as measured, and the validity of individual test score interpretations for the intended uses (AERA et al., 2014). This symposium will present distinct aspects of high-stakes testing that draw from real life applications of tests and are oriented at solving problems that are interesting both from the point of scientific interest and practical use. The first thematic area deals with ideas of equity and access, methodological breadth for fairness, and policy stakes. The symposium interrogates how response formats, scoring policies, and item content interact with gender, language background, and SES to produce measurable DIF and differential outcomes in national and admissions assessments. Using complementary psychometric approaches across distinct European contexts (Sweden and Flanders), the session demonstrates how fairness can be strengthened—or undermined—by design and policy. The second thematic area deals with ideas of validity as a system, decision consequences, and actionability and governance. The session foregrounds validity as an ecosystem, linking interpretive arguments, norm‑referenced data use, and operational safeguards, to improve both the defensibility and usability of examinations internationally. References AERA, APA, & NCME. (2014). Standards for educational and psychological testing. American Educational Research Association. Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73. Messick, S. (1989). Validity. In R. L. Linn (Ed.), Educational measurement (3rd ed., pp. 13–103). Macmillan. Stobart, G. and Eggen, T. (2012). High Stakes testing - value, fairness and consequences. Assessment in Education: Principles, Policy and Practice. Vol.19, pp.1-6. Presentations of the Symposium ***WITHDRAWN*** Standard Setting as an Integral Component of the Validity Argument in High-Stakes Examinations
This paper presents an ongoing study investigating the role of standard-setting in high-stakes examinations for international secondary school exit qualifications. Rather than treating standard-setting as a discrete technical procedure conducted after test construction, the paper advances the position that standard-setting should be conceptualised as an integral component of the overall validity argument for an assessment. Drawing on an argument-based approach to validity (Kane, 2006; Kane, 2013), the study frames standard-setting decisions as evaluative inferences that require explicit warrants, backing, and rebuttals, rather than as neutral psychometric outcomes.
Within this framework, cut scores are understood as consequential judgments that link observed performance to claims about student achievement. These judgments must be supported by coherent evidence relating to construct representation, the quality of expert judgment processes, and the intended and unintended consequences of decisions (Messick, 1989; Mislevy, Almond, & Lukas, 2003). In high-stakes contexts, the defensibility of pass/fail or grade boundary decisions depends not solely on statistical precision, but on the transparency and plausibility of the interpretive and evaluative claims that underpin them (Cizek & Bunch, 2007).
The study is situated within the context of a new national school-leaving qualification which is very high stakes in its context. The resulting of the qualification process, employs a hybrid model of criterion-referenced and norm-referenced standard-setting for its secondary school exit-level examinations. The assessment materials analysed include both fictitious and live examination papers composed of closed-type items, primarily multiple-choice questions and other selected-response formats. These items are designed to sample a defined construct domain aligned to curriculum and qualification specifications.
Methodologically, the study adopts a qualitative-dominant mixed-methods design (Creswell, 2022). Documentary analysis is used to examine the qualification’s standard-setting policies, technical reports, and grade-award documentation. This is complemented by the analysis of standard-setting meeting artefacts, including judges’ performance descriptors, item-level judgments, and rationale statements. Particular attention is paid to how criterion-based expectations are articulated and subsequently moderated through norm-referenced statistical evidence. The analysis maps these processes onto an explicit validity argument, identifying the claims, warrants, and evidence used to justify final cut scores.
The study is currently underway; it is anticipated that hybrid standard-setting models introduce distinctive tensions between construct-based expectations and cohort-relative outcomes. Full findings, including implications for the design and communication of standard-setting decisions in international high-stakes assessment, will be presented at the conference.
References:
Cizek, G. J., & Bunch, M. B. (2007). Standard setting: A guide to establishing and evaluating performance standards on tests. Sage.
Creswell JW. (2022). A concise introduction to mixed methods research (2nd ed.) Sage.
Kane, M. T. (2006). Validation. In R. L. Brennan (Ed.), Educational measurement (4th ed., pp. 17–64). Praeger.
Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73.
Messick, S. (1989). Validity. In R. L. Linn (Ed.), Educational measurement (3rd ed., pp. 13–103). Macmillan.
Mislevy, R. J., Almond, R. G., & Lukas, J. F. (2003). A brief introduction to evidence-centered design. ETS Research Report.
School Leaders’ and Teachers’ Attributions in Making Sense of Norm-referenced School Performance data
School performance data can support school improvement efforts, yet data-informed decision-making is inherently complex and fundamentally interpretative. School leaders and teachers engage in sensemaking by transforming raw data into meaningful information and actionable knowledge. Central to this process are attributions: potential causes professionals construct for school performance outcomes.
Characterising an attribution solely as internal or external is insufficient to meaningfully influence school leaders’ and teachers’ behaviour, as it does not adequately account for the factors that determine their subsequent actions. It is crucial that educational professionals reflect on their role in exercising control over causes for school performance. In response to the lack of a multidimensional view on these attributions in the literature on data use, this study integrates the dimensions of locus of causality (internal versus external) and controllability (controllable versus uncontrollable). Accordingly, we address the following research questions:
RQ1. To which internal/external and controllable/uncontrollable factors do school leaders and teachers attribute school performance on central tests? Do these attributions differ based by work role?
RQ2. Do attributions on central tests vary depending on the type of norm-referenced perspective applied (group referenced versus school referenced)? Do these attributions differ based on the school’s performance level?
Data were collected through 24 semi-structured interviews. A framework analysis was applied, using iterative open and axial coding to identify and refine thematic categories. To enhance reliability, one-third of all coded fragments underwent double coding. To explore trends and differences within and across attributions, categories and perspectives, frequency counts were used to ‘quantitize’ the qualitative data.
Our findings demonstrate that norm-referenced perspectives meaningfully shape the attribution processes of school leaders and teachers. External attributions dominate, although internal factors were recognised as well, particularly at the teaching and school practice levels. Additionally, controllability emerged as a key factor for follow-up actions decisions.
Furthermore, examination of attrbutional differences between population-referenced and group-referenced perspectives reveals that the latter tend to foster internal and external controllable attributions. Thus, the chosen norm-referenced perspective not only directs the locus of causality but also, and perhaps even more importantly, perceptions of controllability and agency.
By adopting this specific focus, this study contributes to the broader discourse on data-informed decision for school improvement both theoretically and practically. It suggests that consciously emphasizing controllable attributions, schools can more effectively translate school performance data into targeted, informed school improvement efforts.
References:
Schildkamp, K. (2019). Data-based decision making for school improvement: Research insights and gaps. Educational Research, 61, 257-273. https://doi.org/10.1080/00131881.2019.1625716.
Feasibility of Using Common AI Models to Improve Rating of Closed-type Items
This paper presents the results of a feasibility study using real paper test items and actual non-trained common LLMs. While in theory (Breuer et al, 2023) closed-type items, such as multiple-choice and similar items, are highly objective and scoring tends to be straightforward, there is still a small proportion of errors. Those come from students‘ ambiguous responses (multiple options selected, crossed out, errors corrected), from unfocused rater or other sources (for example, design flaws of an item). There are several ways to improve the quality and veracity of the rating process and ensure accurate results like better item design, additional checking, etc. (Betts et al, 2021). Since the advance of LLMs and other AI models, there is also an option to harness digital technology to control and reduce the error rate further. This is fundamentally different from the field of automated scoring of open-ended items, where AI models are trained to make estimates about content (Lit et al, 2014). In our case, it is more about using AI for routine and repetitive tasks of transcription or simple scoring.
We will examine four scenarios that combine two different item types (MC item and connecting item) with two rating procedures (rater just transcribing the answer vs. rater awarding points directly). All four scenarios are actual test results from the Slovenian General matura (Grade 13 students) and national assessment (Grade 9 students) for the spring term of 2025. Connection items are from General matura English HE (1869 students) and National Assessment English Gr9 (5165 students), while multiple choice items are from General matura Physics (1429 students) and National Assessment Biology Gr9 (5132 students).
For each scenario, the students‘ responses will be given to AI model with scoring instructions. Whenever the final result (transcribed letter, awarded point) will differ, the response will be examined by a human expert, and his decisions will be final (implicitly correct). This will miss the errors where both AI and humans would score incorrectly, but since the error rates as estimated from appeals and requests for remarking are low to begin with, checking for those would present a disproportionate burden in available resources and wouldn’t be feasible.
The error rates of AI compared to the error rates of human raters will give us the answer if the inclusion of such an additional procedure in a regular testing situation can substantially improve the quality of test results.
References:
Betts, J., Muntean, W., Kim, D., & Kao, S. (2021). Evaluating Different Scoring Methods for Multiple Response Items Providing Partial Credit. Educational and Psychological Measurement, 82, 151 - 176. https://doi.org/10.1177/0013164421994636.
Breuer, S., Scherndl, T., & Ortner, T. (2023). Effects of response format on achievement and aptitude assessment results: multi-level random effects meta-analyses. Royal Society Open Science, 10. https://doi.org/10.1098/rsos.220456.
Liu, O., Brew, C., Blackmore, J., Gerard, L., Madhok, J., & Linn, M. (2014). Automated Scoring of Constructed-Response Science Items: Prospects and Obstacles. Educational Measurement: Issues and Practice, 33, 19-28. https://doi.org/10.1111/emip.12028.
| ||