Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
Please note that all times are shown in the time zone of the conference. The current conference time is: 19th Aug 2026, 21:28:43 EET
|
Daily Overview |
| Session | ||
09 SES 02 A: Fairness and Validity in High-Stakes Testing (Part 1): Fair by Design? Measurement Issues and Equity in High-Stakes Admissions and National Testing
Symposium | ||
| Presentations | ||
09. Assessment, Evaluation, Testing and Measurement
Symposium Fairness and Validity in High-Stakes Testing (Part 1): Fair by Design? Measurement Issues and Equity in High-Stakes Admissions and National Testing The overarching theme is centred around ensuring fairness, validity, and reliability (Kane 2013, Messick, 1989) in high-stakes educational assessments, with a strong emphasis on how different factors (such as item format, gender, language background, socio-economic status, scoring methods, and standard-setting practices) can introduce bias or inequity. High-stakes test are considered “assessments that have serious consequences” for participants (Stobart & Eggen, 2012, p.1) . They typically involve national assessments or external examinations used for admissions to a higher level of education or that otherwise have a significant impact on students’ educational careers and other life choices. With this in mind, fairness and validity are corner concepts of any modern high-stakes testing. Four general aspects of fairness are the equivalent opportunity of test-takers in the testing process, the lack of measurement bias, the access to the construct(s) as measured, and the validity of individual test score interpretations for the intended uses (AERA et al., 2014). This symposium will present distinct aspects of high-stakes testing that draw from real life applications of tests and are oriented at solving problems that are interesting both from the point of scientific interest and practical use. The first thematic area deals with ideas of equity and access, methodological breadth for fairness, and policy stakes. The symposium interrogates how response formats, scoring policies, and item content interact with gender, language background, and SES to produce measurable DIF and differential outcomes in national and admissions assessments. Using complementary psychometric approaches across distinct European contexts (Sweden and Flanders), the session demonstrates how fairness can be strengthened—or undermined—by design and policy. The second thematic area deals with ideas of validity as a system, decision consequences, and actionability and governance. The session foregrounds validity as an ecosystem, linking interpretive arguments, norm‑referenced data use, and operational safeguards, to improve both the defensibility and usability of examinations internationally. References AERA, APA, & NCME. (2014). Standards for educational and psychological testing. American Educational Research Association. Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73. Messick, S. (1989). Validity. In R. L. Linn (Ed.), Educational measurement (3rd ed., pp. 13–103). Macmillan. Stobart, G. and Eggen, T. (2012). High Stakes testing - value, fairness and consequences. Assessment in Education: Principles, Policy and Practice. Vol.19, pp.1-6 Presentations of the Symposium DIF across Genders and Swedish Language Syllabus Groups in the Swedish National Test in Social Studies: The Role of Response Formats
In developing the national tests in Social Studies, a critical issue is to ensure that the items work in the same way for different classifications of students, i.e. that the test fair and show no signs of severe Differential Item Functioning (DIF). Social Studies is a mandatory subject in the Swedish compulsory school, covering a wider range of content than, for instance, the subject of Civics. In Sweden, the national tests are intended to be digitalized and the proportions of items with selected response format are also intended to increase. As a consequence, the proportion of items with a constructed response format will decrease. It is important analyse the consequences of this shift for different subgroups of students. However, preliminary analyses (Löfstedt & Bergh, 2025) show that there are indeed DIF across genders with respect to different item formats, where boys commonly find selected response item formats easier compared to girls. In addition, in a forthcoming article (Löfstedt & Bergh, 2026) it is concluded that there is a risk that students following the syllabus of Swedish as a second language to be disadvantaged compared to the general student population, if the proportion of selected response items is increased. The purpose of this study is to analyse DIF across genders and Swedish language syllabus groups, taking into account potential subgroup interactions. The study is based on data from the Swedish national test in Social Studies collected in 2018, 2019, 2022 and 2023. The data are analysed by means of the Polytomous Rasch Model (Rasch, 1960/1980; Andrich, 1978). The results reveal that there are indeed important DIF effects across genders and Swedish language syllabus groups, having significant implications for subgroups of students. Thus, as the test result influences the student’s final subject grade, the results of this study may have implications for student opportunities for upper secondary school program admission. However, the patterns are clarified by also taking into account the subgroup interactions (gender by Swedish language syllabus) in the DIF analyses. In addition, consequences of the item format shift are discussed from a test dimensionality perspective.
References:
Andrich, D. (1978). A rating formulation for ordered response categories. Psychometrica, 43 (4), 561-573.
Löfstedt, A., & Bergh, D. (2025, September 2–3). The question of gender in national testing: Potential implications of format change as a consequence of digitalization for subgroups of students—Experiences from the Swedish national test in social studies [Conference presentation]. Frontier Research in Educational Measurement (FREMO), Oslo, Norway.
Löfstedt, A., & Bergh, D. (2026, in press). Investigating DIF: The role of response format for students following the syllabus for Swedish as a second Language – Experiences from the Swedish National Tests in Social Studies. Nordidactica 2025:2.
Rasch, G. (1960/1980). Probabilistic models for some intelligence and attainment tests (Copenhagen, Danish Institute for Educational Research). Expanded edition (1980) with foreword and afterward by Benjamin D. Wright. Chicago: The University of Chicago Press.
Exploring Gender Differences in Students’ Response Behavior on the Flemish Medical Entrance Exam
In Flanders, matriculation to medical education is made dependent upon a centrally organized admission test assessing competences in sciences (mathematics, physics, chemistry, and biology), interpersonal skills, and scientific text comprehension using multiple-choice items that are scored using the so-called “correction for guessing”. Admission is granted to a fixed number of examinees who are ranked highest among those examinees that obtained a passing score on the different subtests.
In 2023, the acceptance rate equaled .53 for male examinees who just finished high school while being only .38 for their female peers. Similar gender differences were found among the group of older examinees and examinees coming from abroad. Two factors can explain these gender differences. Firstly, due to self-selection male applicants may on average be stronger in the assessed skills than female applicants. Secondly, awarding a negative score to wrong answers may incentive participants to omit response to items when they do not know or are not sure of the answer, which could impact women more than men (e.g., Kelly & Dennick, 2009).
To disentangle both factors, a tree-based Item Response Theory model (De Beer, Janssen, De Boeck, 2017) is applied to the results of the 2021, 2022, and 2023 science tests of the Flemish medical entrance examination. By modeling the correctness of given responses and response omission jointly, a separate estimate of examinees’ science proficiency and skipping propensity is obtained. Gender differences are observed on both measures. Moreover, by simulating responses on the skipped items based on examinees’ science proficiency, the effect of the gender differences in skipping propensity on the ranking of examinees is illustrated.
Hence, the present study shows that both factors contribute to the gender difference in acceptance rates for all three investigated years of administration. The study therefore argues to monitor the quality of high-stakes examination in several respects in order to avoid gender bias, especially with numerus fixus (ranking of candidates). Moreover, exam committees should consider alternative instructions for negative marking in high-stakes multiple-choice exams, such as elimination scoring in which students can show their partial knowledge on exam questions (Vanderoost et al., 2019).
References:
De Beer, D., Janssen, R., & De Boeck, P. (2017). Modeling skipped and notreached items using IRTrees. Journal of Educational Measurement, 54(3), 333–363. https://doi.org/10.1111/jedm.12147
Kelly, S., & Dennick, R. (2009). Evidence of gender bias in True-False-Abstain medical examinations. BMC Medical Education, 7, 9-32. doi: 10.1186/1472-6920-9-32.
Vanderoost, J., Janssen, R., Callens, R., Eggermont, J., De Laet, T. with De Laet, T. (corresp. author) (2018). Elimination testing with adapted scoring reduces guessing and anxiety in multiple-choice assessments, but does not increase grade average with respect to traditional negative marking. PLoS One, 13 (10). doi: 10.1371/journal.pone.0203931
Internal Structure Validity Evidence for the SweSAT, based on Test Takers’ Socio-economic Status and Language Background
Gathering evidence for or against the validity argument of large-scale tests is a crucial aspect of maintaining their integrity. One important source of such evidence is the test’s internal structure, which can be investigated using Differential Item Functioning (DIF) analyses. The present study is a comprehensive investigation of DIF in the Swedish Scholastic Aptitude Test (SweSAT), a large-scale test for selection to higher education, related to test takers’ socio-economic status (SES) and language background.
In this study, the test takers’ highest parental education level was used as the indicator of SES. Language background was operationalized as whether a test taker took the course “Swedish as a second language” in elementary schoolyear 9 or not. Data on SES and language background were collected when an individual entered the Evaluation Through Follow-up (UGU) dataset at age 16. SweSAT scores were collected later. Using the Mantel-Haenszel procedure on archival data from approximately 3,360 items and nearly one million test takers, the study examined whether any test items systematically favored any group.
We hypothesized that DIF would be minimal in the quantitative section and more pronounced in the verbal section, mirroring patterns and their proposed explanations observed previously with gender as the grouping variable. In the verbal domain, vocabulary items were expected to contain the most DIF, followed by sentence completion items and English reading comprehension items, with Swedish reading comprehension items containing the least DIF. Vocabulary items potentially favored groups more familiar with the word’s contextual domain (for example, words drawn from everyday life might have favored lower-SES test takers and those with Swedish as a second language, while words requiring specialized cultural knowledge might have disfavored those groups). Other than that, no specific patterns were expected.
The analysis is still in progress, and the results will be presented at the conference. Overall, the findings are expected to contribute to the SweSAT and the public by adding validity evidence based on SES and language background for or against the SweSAT’s validity argument. Any identified DIF patterns attributable to bias will be used to formulate recommendations for enhancing test fairness, such as refining item content or format, thus further strengthening the test’s validity argument and transparency.
References:
Allalouf, A., Hambleton, R. K., & Sireci, S. G. (1999). Identifying the causes of DIF in translated verbal items. Journal of Educational Measurement, 36(3), 185–198.
Carlton, S. T., & Harris, A. M. (1992). Characteristics associated with differential item functioning on the Scholastic Aptitude Test: Gender and majority/minority group comparisons. (Report no. ETS-RR-92-64). Princeton, NJ: Educational Testing Service.
Holland, P. W., & Thayer, D. T. (1988). Differential item performance and the Mantel- Haenszel procedure. In H. Wainer & H. I. Braun (Eds.), Test validity (pp. 129– 145). Hillsdale NJ: Erlbaum.
Schmitt, A. P., Holland, P. W., & Dorans, N. J. (1992). Evaluating hypotheses about differential item functioning. (Report No. ETS-RR-92-8). Princeton, NJ: Educational Testing Service.
Stage, C. (2005). Socialgruppsskillnader i resultat på högskoleprovet [Social group differences in scores on the Swedish Scholastic Assessment Test] (Report No. BVM 11:2005). Umeå, Sweden: Umeå University.
Wedman, J. (2017). Reasons for gender-related differential item functioning in a college admissions test, Scandinavian Journal of Educational Measurement, 62(6), 959–970. DOI: 10.1080/00313831.2017.1402365
Wester, A. (1997). Differential item functioning (DIF) in relation to item content: A study of three subtests in the SweSAT with focus on gender. (Report No. EM No 27). Umeå, Sweden: Umeå University.
| ||
