Conference Agenda
| Session | ||
09 SES 12 C: Psychometric Modelling, Adaptive Testing and Item Complexity
Paper Session | ||
| Presentations | ||
09. Assessment, Evaluation, Testing and Measurement
Paper Developing a Live Cognitive Diagnostic Multistage Test for Cross-Cultural Assessment of Fourth-Grade Mathematics Skills in Türkiye and Spain 1: Bogazici University, Turkey (Türkiye); 2: Comillas Pontifical University, Spain; 3: Pamukkale University, Turkey (Türkiye); 4: The University of Hong Kong, China Presenting Author:The primary purpose of educational assessment is to support learning by identifying students’ strengths and weaknesses rather than merely reporting a single performance score. Nevertheless, summative assessment practices—typically based on Classical Test Theory or Item Response Theory (IRT)—remain dominant in many educational systems (Hambleton et al., 1991). Although linear tests might provide reliable overall ability estimates, they offer limited diagnostic information about learners’ mastery of specific skills (Choi, 2010; de la Torre & Karelitz, 2009). Cognitive Diagnostic Models (CDMs) have been developed to address this limitation by classifying examinees according to mastery or non-mastery of predefined skills or attributes (de la Torre & Douglas, 2004). CDMs enable the provision of structured, fine-grained feedback that can support formative assessment and instructional decision-making, particularly in primary education where early identification of learning gaps is crucial (Huebner, 2010). In parallel, advances in digital testing have led to increasing use of adaptive testing approaches such as Computerized Adaptive Testing (CAT) and Multistage Testing (MST). Adaptive tests can improve measurement precision while reducing test length (Embretson, 1996; Weiss, 2004). Among these approaches, MST offers practical advantages for educational settings, including content balancing, quality control, and examinee navigation flexibility (Sari et al., 2016; Yan et al., 2014). However, MST implementations are typically based on IRT and therefore still yield a single overall score. To enhance the formative potential of adaptive assessment, recent research has focused on integrating CDMs with adaptive testing. Cognitive Diagnostic Computerized Adaptive Testing (CD-CAT) and Cognitive Diagnostic Multistage Testing (CD-MST) represent promising approaches for delivering mastery-based feedback within adaptive frameworks (Cheng, 2009; de la Torre, 2019; Kim & Yoo, 2023; von Davier & Cheng, 2014). Despite this promise, empirical research on CD-MST remains limited and has largely relied on simulation studies (Kim & Yoo, 2023; Li et al., 2021; Liu et al., 2018). Evidence from live, large-scale implementations—especially with young learners—is scarce. Beyond methodological innovation, the use of a live CD-MST is particularly important for classroom assessment because it bridges the gap between psychometric sophistication and instructional usability. Unlike simulation-based studies, live tests capture authentic student behaviors, contextual constraints, and instructional realities that directly affect the validity and utility of diagnostic feedback. Collecting live data therefore allows for evaluation of whether CD-MST results are interpretable, stable, and actionable for teachers and learners in real educational settings. Moreover, empirical evidence from classroom use is essential to demonstrate CD-MST’s practical feasibility for formative assessment and inform educators and policymakers about its potential to support data-driven instruction and early intervention—particularly in primary education, where timely and accurate diagnostic information can have long-lasting effects on students’ learning trajectories. Thus, the present study addresses this gap by developing and implementing a live CD-MST for fourth-grade mathematics in Türkiye and Spain. The test is designed to provide diagnostic feedback on core arithmetic skills (addition, subtraction, multiplication, division, routine and nonroutine problem solving) while ensuring cross-cultural comparability. By administering the test digitally in real school settings, the study aim to contribute empirical evidence on the feasibility, validity, and formative value of CD-MST in European educational contexts. The guiding research question is: To what extent do Turkish and Spanish fourth-grade students demonstrate mastery of mathematics skills as assessed by a live Cognitive Diagnostic Multistage Test? Methodology, Methods, Research Instruments or Sources Used Participants The study will be conducted in two phases. During the pilot phase, data will be collected from approximately 800 fourth-grade students (400 in Türkiye and 400 in Spain). In the live CD-MST phase, an additional 200 students (100 per country) will participate, resulting in a total sample of 1000 students. Due to school-level permission requirements, convenience sampling will be employed. In Türkiye, participating schools will be recruited from the Istanbul region, and in Spain from the Madrid region. Efforts will be made to include schools from diverse socioeconomic backgrounds and both public and private sectors to ensure variability in student ability levels. Students will be aged 9–10 years, with an approximately balanced gender distribution. All testing will take place digitally (computers or tablets) within school environments, typically in computer laboratories. Each session will last approximately 50 minutes. Participation will be voluntary, with no incentives provided. Ethical approvals and relevant institutional permissions were obtained prior to data collection. Instrument A common table of specifications aligned with both Turkish and Spanish mathematics curricula were developed. From an existing pool of 540 fourth-grade mathematics items aligned with TIMSS cognitive domains, 60 items were selected based on curricular relevance, difficulty range, and suitability for CD-MST. Item adaptation from Turkish to Spanish followed a multistage translation procedure, in which the original items were first translated into English and then translated into Spanish by experts to ensure linguistic and conceptual equivalence. Validity evidence will be gathered using Q-matrix validation procedures. The pilot phase for estimating item parameters will employ a multiple-matrix sampling design with four linked booklets, each containing 20 items, to reduce students’ testing burden. Additionally, the differential item functioning (DIF) analyses will be conducted and biased items will be excluded before final test assembly. Development of CD-MST The final live CD-MST will employ a 1–3–3 multistage design, with approximately 12–15 items administered per student. Both the pilot and final live assessments will be implemented using the Concerto platform. Module selection will be guided by indices based on the Kullback–Leibler distance or the generalized DINA model discrimination index. Data Analysis After development, a sample of students (100 Turkish and 100 Spanish) will take the CD-MST to assess its functionality. Students’ mastery profiles will be estimated using DINA or G-DINA models (de la Torre, 2011). Attribute-level probabilities will be reported to address the research question. Conclusions, Expected Outcomes or Findings The study is expected to demonstrate the feasibility of implementing a live CD-MST with primary school students in authentic classroom settings. The results are anticipated to show that diagnostic information can be obtained with a substantially reduced testing burden compared to traditional fixed-form assessments. Previous simulation studies have indicated that CD-CAT achieves classification accuracy comparable to that of paper-and-pencil CDM tests (de la Torre, 2019); therefore, similar levels of accuracy with CD-MST across cultures are expected. From a European perspective, findings are expected to provide empirical evidence on the cross-cultural applicability of diagnostic adaptive assessments. Mastery profiles are expected to reveal meaningful patterns in students’ strengths and weaknesses, highlighting the potential of CD-MST to support formative assessment and early intervention across educational systems. Methodologically, the study will contribute to the limited body of empirical research on live CD-MST implementations and inform future European initiatives in digital assessment design. Practically, the findings may support the broader adoption of mastery-based adaptive testing as a tool for promoting equity by providing all learners with access to individualized, data-driven feedback regardless of resource constraints. References Cheng, Y. (2009). When cognitive diagnosis meets computerized adaptive testing: CD-CAT. Psychometrika, 74, 619-632. Choi, H. J. (2010). A model that combines diagnostic classification assessment with mixture item response theory models (Doctoral dissertation). University of Georgia, Athens de la Torre, J. (2019, June). Recent advances in cognitive diagnosis computerized adaptive testing. Paper presented at the Seventh International Meeting of the International Association for Computerized Adaptive Testing, Minneapolis. de La Torre, J., & Douglas, J.A. (2004). Higher-order latent trait models for cognitive diagnosis. Psychometrika 69, 333–353 de La Torre, J., & Karelitz, T. M. (2009). Impact of diagnosticity on the adequacy of models for cognitive diagnosis under a linear attribute structure: A simulation study. Journal of Educational Measurement, 46(4), 450-469 Embretson, S. E. (1996). The new rules of measurement. Psychological Assessment, 8(4), 341-349 Hambleton, R. K., Swaminathan, H., & Rogers, H. J. (1991). Fundamentals of item response theory (Vol. 2). Sage. Huebner, A. (2010). An overview of recent developments in cognitive diagnostic computer adaptive assessments. Practical Assessment, Research, and Evaluation, 15(3), 1-7. Kim, R. Y., & Yoo, Y. J. (2023). Cognitive diagnostic multistage testing by partitioning hierarchically structured attributes. Journal of Educational Measurement, 60(1), 126-147. Li, G., Cai, Y., Gao, X., Wang, D., & Tu, D. (2021). Automated test assembly for multistage testing with cognitive diagnosis. Frontiers in Psychology, 12, 509844. Liu, S., Cai, Y., & Tu, D. (2018). On-the-Fly Constraint-Controlled Assembly Methods for Multistage Adaptive Testing for Cognitive Diagnosis. Journal of Educational Measurement, 55(4), 595-613. Sari, H. İ., Yahsi-Sari, H., & Huggins-Manley, A. C. (2016). Computer adaptive multistage testing: Practical issues, challenges and principles. Journal of Measurement and Evaluation in Education and Psychology, 7(2), 388-406. Weiss, D. J. (2004). Computerized Adaptive Testing for Effective and Efficient Measurement in Counseling and Education. Measurement and Evaluation in Counseling and Development, 37(2), 70–84. von Davier, M. & Cheng, Y. (2014). Multistage testing using diagnostic models. In D. L. Yan, A. A. von Davier, & C. Lewis (Eds.), Computerized multistage testing: Theory and practice (pp. 219–227). Boca Raton, FL: Chapman & Hall/CRC. Yan, D., Lewis, C., & von Davier, A. A. (2014). Overview of computerized multistage tests. In D. Yan, C. Lewis, & A. A. von Davier (Eds.), Computerized multistage testing: Theory and applications (pp. 3–20). New York: Chapman and Hall/CRC. 09. Assessment, Evaluation, Testing and Measurement
Paper Improving Performance Estimation Under Ceiling Effects in Adaptive Testing: A Simulation and Empirical Study Using Plausible Values 1: Finnish Education Evaluation Centre, Finland; 2: Tampere University, Finland; 3: University of Helsinki, Finland Presenting Author:Computerised adaptive testing (CAT) is increasingly used in large-scale educational assessment studies. Compared to paper-based fixed tests, CAT has several advantages, including increased accuracy of the measurement and a better match of the difficulty level of the test with the test taker’s abilities (van der Linden & Glas, 2000). The idea of adapting the test’s difficulty to the test taker originates from individually administered traditional intelligence tests (e.g., Wechsler, 1944), but computer-based testing has made it possible to administer the tests also in group settings. The main idea of CAT is to have a validated item bank, in which each item has difficulty and discrimination parameters that have been defined in advance based on calibration data using IRT (item response theory) modeling (Birnbaum, 1968). After a fixed set of anchor items, the subsequent items are selected from the item bank based on the test taker’s performance. Even though CAT solves many measurement issues compared to a more traditional approach with fixed tests, accurate assessment of the abilities of test takers who demonstrate significantly lower or higher performance than average can still be challenging. In educational assessment studies utilising item banks calibrated with data collected in school contexts, these problems may emerge when trying to assess the performance level of particularly highly skilled students or lower-performing students. Despite the large calibration data and the number of items, the performance estimates may not be trustworthy if a student does not fail enough items, or if they fail all of them within the limited time allowed for completing the test. In these occasions, adaptive testing does not solve the problems often faced in group-level assessments utilising fixed tests, including the so-called ceiling effect (Wang et al., 2009) which can distort analyses where the ability distributions of different sub-groups are compared on the population level. Thus, the IRT-based approach alone does not provide sufficient information on student performance if the testing time is not extended or if the difficulty level of the starting point of the test is not individually manipulated. This is usually not possible in large-scale group-level assessments. Therefore, in the context of a large-scale longitudinal assessment study with a particular focus on students with special educational needs, we aim at solving this measurement problem by – besides considering students’ performance – taking into account multiple background-related factors associated with performance in earlier assessment studies. We first conducted a simulation study to compare different performance estimation methods in a longitudinal context (IRT, plausible values) and then applied the most appropriate solution on the empirical data. Methodology, Methods, Research Instruments or Sources Used In the simulation study, we explored whether plausible values can better capture the true differences between subgroups in the population when there are severe ceiling effects in the raw test scores or IRT estimates. Under two conditions (ceiling effects vs. no ceiling effects), we created six different scenarios that differed from each other in how subgroup differences evolved between three time points. The empirical analysis of the study will be based on the data (N ~ 3100) from the longitudinal study conducted in Finnish primary schools (ISCED level 1). In this single-cohort study, the same students were followed yearly from the 4th grade to the 6th grade in the spring semesters of 2022–2024. The analysis focuses on students’ arithmetical reasoning skills that were assessed with tasks that adjusted to each student’s performance level. The data come from a computerised adaptive test that was calibrated with over 10000 primary and secondary school students in Finland. After four anchor items that were the same for all students, the task set presented the student with either slightly more difficult or easier items depending on the student’s responses. The task set was programmed to end automatically once the student had completed 15 tasks or had spent 15 minutes on the test. In the empirical study, we will create plausible values for the students based on their raw performance scores and their background (gender, parents’ occupational status, immigration background, special educational needs). We will study whether the use of plausible values will influence the conclusions made on student performance and its development compared to the IRT-based scoring (ML type scores). We focus particularly on student subgroups whose performance deviates considerably from average students for whom the adaptive test is primarily calibrated. Conclusions, Expected Outcomes or Findings The results of the preliminary simulations indicate that the use of plausible values can to some extent correct the ceiling effects observed when the estimation of performance is only based on IRT. In the next stage of the analysis, we will create plausible values for students in the empirical data and conduct corresponding comparisons. Based on the results of the preliminary simulations, we expect to increase the accuracy of the measurement of performance particularly for highly skilled students who reached the end of the testing time before finding the upper limit of their abilities. From the perspective of the research project providing the empirical data for the study, it is also important to see whether we manage to improve the accuracy of estimation for students with special educational needs for whom the group-level adaptive test starts with items that are too difficult. The results of the study can be used in the design of assessment projects focusing on educational equality, as such projects often analyse the results of student subgroups whose performance deviates from that of the average student at a certain age or grade level. References Birnbaum, A. (1968). Some latent trait models and their use in inferring an examinee’s ability. In F. M. Lord & M. R. Novick (Eds.), Statistical theories of mental test scores (pp. 397–479). Addison-Wesley. van der Linden, W. J. & Glas, C. A. W. (2000). Computerized adaptive testing : theory and practice. Kluwer Academic Publishers. https://doi.org/10.1007/0-306-47531-6 Wang, L., Zhang, Z., McArdle, J. J., & Salthouse, T. A. (2009). Investigating ceiling effects in longitudinal data analysis. Multivariate Behavioral Research, 43(3), 476–496. https://doi.org/10.1080/00273170802285941 Wechsler, D. (1944). The measurement of adult intelligence (3rd ed.). The Williams & Wilkins Company. 09. Assessment, Evaluation, Testing and Measurement
Paper Factor Mixture Modeling of Confidence in Science: Examining Latent Class–Specific Measurement Structures in TIMSS 2023 1: Canakkale Onsekiz Mart University, Türkiye; 2: Aydin Adnan Menderes University, Türkiye Presenting Author:Confidence in science is a central affective construct that has been consistently linked to students’ engagement, persistence, and achievement in science learning (Berger et al., 2025; Lee & Stankov, 2018; Wu & Wu, 2022). Students with higher confidence in science are more likely to participate actively in learning activities (OECD, 2019) and to pursue further studies in science-related fields (Wang & Degol, 2013). Recently released results from TIMSS 2023 provide further evidence that eighth-grade students who report being very confident in science demonstrate substantially higher average science achievement (von Davier et al., 2024). Consequently, understanding how confidence in science is conceptualised and distributed among students is of critical importance for educational research. In the TIMSS 2023 Assessment Frameworks (Reynolds et al., 2021), confidence in science is conceptualised under the students’ attitudes toward science. It is commonly measured using self-report Likert-type scales. The scale scores are reported using three categorical levels (e.g., not confident, somewhat confident, and very confident) defined through cut scores that are determined using Latent Class Analysis (LCA; Lazarsfeld & Henry, 1968). While the LCA-based approach adopted in TIMSS offers a data-driven method for identifying qualitatively distinct groups of students, it conceptualises confidence in science, assuming measurement homogeneity within each identified group. This assumption may be limiting in large-scale assessments, where latent subgroups may differ in overall levels of confidence and in how the construct is structured and measured. Factor Mixture Modeling (FMM) provides an integrative framework that combines confirmatory factor analysis (CFA) and latent class analysis (LCA), allowing researchers to simultaneously examine categorical heterogeneity and within-class variation (Muthén, 2006). Importantly, FMM provides a flexible approach to investigate potential differences in the latent construct and measurement properties of confidence in science across unobserved subgroups, providing a means to investigate measurement equivalence. The present study focuses on a comparison between Türkiye and England, two countries that differ substantially in their educational systems, instructional traditions, and sociocultural contexts (OECD, 2023). According to TIMSS 2023 results, the average science achievement scores for Grade 8 are highly similar in the two countries (von Davier et al., 2024), making this comparison potentially informative for examining whether comparable performance levels reflect comparable latent constructs. Using FMM, this study examines whether the same number of latent classes emerges in both countries and whether the measurement structure of confidence in science operates equivalently across these latent classes. Addressing these issues is essential for evaluating measurement equivalence and for informing the comparability of the scale scores across countries. Accordingly, the study addresses the following research questions:
Methodology, Methods, Research Instruments or Sources Used To address the research questions, data were retrieved from the TIMSS 2023 international database and drawn from the Grade 8 student context questionnaire administered in Türkiye and England. The Turkish sample comprised 4,925 students initially, with 4,508 retained after data cleaning. The sample comprised of 50.4% girls (n = 2271), and 49.6% boys (n = 2237). The mean age of students was 13.87 years (SD = 0.41). The English sample comprised 4,239 students initially, with 4,190 retained after data cleaning. The sample comprised of 51.4% girls (n = 2155), and 48.6% boys (n = 2035). The mean age of students was 14.04 years (SD = 0.30). The TIMSS 2023 Students Confident in Science scale captures students’ self-perceptions of their ability to learn and succeed in science, which are informed by prior experiences and comparative self-evaluations (Reynolds et al., 2021). These items were scored on a 4-point Likert-type scale, ranging from 1 = agree a lot to 4 = disagree a lot. At this stage, preliminary analyses and FMMs were first conducted for the Turkish sample to inform model specification. Corresponding analyses are planned for the English sample. The analytical procedure followed a stepwise modeling strategy (Clark et al., 2013). Baseline models were evaluated using CFA to assess unidimensionality and LCA to identify latent classes. In the most restrictive specification (FMM-1), factor loadings and item thresholds were constrained to be invariant across latent classes. Subsequent models (FMM-2 to FMM-4) progressively relaxed these constraints by allowing class-specific factor variances, thresholds, and loadings, with factor means fixed for model identification. Based on these preliminary findings, a sequence of increasingly flexible FMMs will subsequently be evaluated within a multigroup framework. Measurement invariance across latent classes will be examined by treating country as a known grouping variable, provided that an acceptable baseline model is established in each country. Acceptable model fit for CFA was indicated by RMSEA and SRMR values below 0.08, together with CFI and TLI values exceeding 0.90 (Hu & Bentler, 1999). For LCA and FMM, model selection was informed by the BLRT, LMR, and VLMR tests, where significant p-values indicate improved fit for models with additional classes (Nylund et al., 2007). Information criteria (AIC, the BIC, and the SABIC) were also examined, with lower values indicating better relative fit. Classification quality was evaluated using entropy, with values closer to 1 indicating greater classification precision. Conclusions, Expected Outcomes or Findings At this stage, conclusions are based on the Turkish sample; analyses involving the English sample and multigroup FMM will be undertaken in a subsequent phase of the study. In the Turkish sample, one- and two-factor models was evaluated to examine factor structure. CFA results indicated that both models demonstrated acceptable fit; however, the one-factor model was preferred due to its more parsimonious structure and comparable fit indices (CFI = .99, TLI = .99, RMSEA = .064, SRMR = .016). LCA findings indicated that a three-class solution provides a reasonable representation of the underlying heterogeneity. Preliminary FMM results indicate systematic improvements in model fit as measurement constraints were progressively relaxed. In both the two- and three-class solutions, more flexible specifications allowing measurement noninvariance (FMM-3 and FMM-4) yielded lower information criteria and higher classification quality. This pattern suggests that the observed differences between classes are not solely attributable to differences in the underlying factor (Lubke & Muthén, 2005), but may also reflect class-specific measurement characteristics. Accordingly, differences between the latent classes may reflect qualitative distinctions rather than purely quantitative differences. These findings highlight the potential utility of FMM for examining class-specific measurement structures within heterogeneous samples. References Berger, N., Fitzmaurice, O., Ryan, V., Mackenzie, E., & Holmes, K. (2025). Secondary students’ attitudes towards mathematics and science in Ireland: a latent profile analysis of TIMSS 2019 using situated expectancy-value theory. Irish Educational Studies, 1-23. https://doi.org/10.1080/03323315.2025.2549098 Clark, S. L., Muthén, B., Kaprio, J., D'Onofrio, B. M., Viken, R., & Rose, R. J. (2013). models and strategies for factor mixture analysis: an example concerning the structure underlying psychological disorders. Structural Equation Modeling: A Multidisciplinary Journal, 20(4), 681-703. http://dx.doi.org/10.1080/10705511.2013.824786 Lazarsfeld, P. F., & Henry, N. W. (1968). Latent structure analysis. New York: Houghton Mifflin. Lee, J., & Stankov, L. (2018). Non-cognitive predictors of academic achievement: Evidence from TIMSS and PISA. Learning and Individual Differences, 65, 50-64. https://doi.org/10.1016/j.lindif.2018.05.009 Lubke, G. H., & Muthén, B. (2005). Investigating population heterogeneity with factor mixture models. Psychological Methods, 10(1), 21-39. https://doi.org/10.1037/1082-989X.10.1.21 Muthén, B. (2006). Should substance use disorders be considered as categorical or dimensional? . Addiction, 101(1), 6-16. https://doi.org/10.1111/j.1360-0443.2006.01583.x Nylund, K. L., Asparouhov, T., & Muthén, B. O. (2007). Deciding on the number of classes in latent class analysis and growth mixture modeling: a monte carlo simulation study. Structural Equation Modeling: A Multidisciplinary Journal, 14(4), 535-569. https:// doi.org/ 10. 1080/ 10705 51070 1575396 OECD. (2019). PISA 2018 Results (Volume II): Where All Students Can Succeed. Paris: OECD Publishing. https://doi.org/10.1787/b5fd1b8f-en OECD (2023). Education at a Glance 2023: OECD Indicators. Paris: OECD Publishing. https://doi.org/10.1787/e13bef63-en. Reynolds, K. A., Mullis, I. V., & Martin, M. O. (2021). Chapter 3: TIMSS 2023 Context Questionnaire Framework. In I. V. Mullis, M. O. Martin, & M. von Davier, TIMSS 2023 Assessment Frameworks (s. 46-71). TIMSS & PIRLS International Study Center: Boston College. https://doi.org/10.6017/lse.tpisc.timss.rs6460 von Davier, M., Kennedy, A., Reynolds, K., Fishbein, B., Khorramdel, L., Aldrich, C., . . . Yin, L. (2024). TIMSS 2023 International Results in Mathematics and Science. Boston College: TIMSS & PIRLS International Study Center. https://doi.org/10.6017/lse.tpisc.timss.rs6460 Wang, M. T., & Degol, J. (2013). Motivational pathways to STEM career choices: Using expectancy–value perspective to understand individual and gender differences in STEM fields. Developmental Review, 33(4), 304-340. https://doi.org/10.1016/j.dr.2013.08.001 Wu, X., & Wu, Q. (2022). Can students be more engaged and confident? A multiple membership multilevel analysis of science engagement and confidence and their effects on science achievement in East Asia. Studies in Educational Evaluation, 73(101147). https://doi.org/10.1016/j.stueduc.2022.101147 09. Assessment, Evaluation, Testing and Measurement
Paper Modeling the Interaction Between Person-Level Covariates and Item Complexity in Creative Thinking Task Performances in PISA2022 Trakya University, Turkey (Türkiye) Presenting Author:Language background plays a central role in how students understand, process, and respond to items in cognitive assessments. Extensive research indicates that even subtle changes in test item wording can significantly influence student performance (Abedi et al., 2000; Abedi & Lord, 2001). For instance, employing simpler vocabulary, shorter sentences, and active voice allows students to demonstrate their true abilities more accurately. By reducing the linguistic burden, educators ensure that assessments measure intended knowledge rather than a student's ability to decode complex syntax. This issue is particularly critical in international and European assessment contexts. In many countries, students participate in large-scale assessments conducted in a language that is not their first. While these students may possess profound knowledge in subjects like mathematics or science, they often perform poorly because they struggle with the specific language used in test items. Evidence suggests these language barriers reduce both the validity and reliability of test scores, creating systemic disadvantages for English Language Learners (ELLs) and students from language-minority backgrounds (Abedi, 2002, 2004). From a measurement perspective, this phenomenon is explained by construct-irrelevant variance (Messick, 1994). This occurs when test scores are influenced by factors that are not part of the specific construct intended for measurement. In this context, linguistic complexity acts as an unintended source of difficulty. Consequently, test scores may inadvertently reflect a student’s reading proficiency or language exposure rather than their actual mastery of the subject matter. Previous research has pinpointed several linguistic features that negatively affect comprehension. Long, convoluted sentences, unfamiliar academic vocabulary, and passive voice increase cognitive load, slowing down reading and making misunderstandings more likely (Abedi, 2006). While math performance typically improves when wording is concise, most existing studies focus on traditional domains. There is currently a significant gap regarding Creative Thinking Tasks (CTTs). This is a crucial oversight, as CTTs frequently utilize open-ended and language-rich prompts. Students must generate original ideas and justify their reasoning, requiring not only divergent thinking but also high-level language comprehension. Because of this, item wording may exert an even stronger influence on CTTs than on standardized multiple-choice tests. Students with high creative potential but limited language exposure may be severely disadvantaged, raising concerns regarding fairness and equity in international assessments. Recent evidence from ePIRLS 2016 (Bulut et al., 2023) supports this, showing that item complexity significantly reduces the probability of success in online reading. This study aims to examine if similar effects occur in creative thinking tasks. By considering both item-level factors (complexity) and student-level factors (reading achievement, language background, gender, and socioeconomic status), the study contributes to the global dialogue on assessment validity.
Methodology, Methods, Research Instruments or Sources Used The factors influencing students’ probability of correctly answering PISA 2022 CTT items were examined using a two-level Generalized Linear Mixed Model (GLMM) approach. The dataset comprised responses from 83,295 students across 15 countries, all of whom completed the cognitive assessment in English. For the purposes of analysis, Level 1 (item level) included the linguistic complexity of the items, while Level 2 (student level) included reading achievement, language most frequently spoken at home (test language [English] vs. another language), time spent on homework related to the test language (English) per day, socioeconomic status (ESCS), and gender. Item linguistic complexity was operationalized using ATOS readability scores for the 18 released CTT items (Renaissance Learning, 2012). Higher ATOS scores indicate greater linguistic complexity. The ATOS scores were computed based on the item texts publicly available at https://pisa2022-questions.oecd.org/. ATOS values ranged from 5.37 to 9.14 (M = 7.45, SD = 1.02). The least linguistically complex item was Space Comic Q02, whereas Carpooling Q01 exhibited the highest linguistic complexity. At the student level, analyses were repeated for each of the ten plausible values of reading achievement, and the results were pooled according to Rubin’s rules (Rubin, 1987). Students with missing responses on the language background, homework, ESCS, or gender variables were excluded from the analyses, resulting in a final analytic sample of 73,135 students. The probability of correctly answering CTT items was examined through the relative comparison of four hierarchical models: • Model 0: Null model • Model 1: Item-level predictors only • Model 2: Model 1 + student-level predictors • Model 3: Model 2 + moderation effects of student-level predictors Model comparisons were conducted using the Akaike Information Criterion (AIC), Bayesian Information Criterion (BIC), log-likelihood (LL), and R² indices (Nakagawa & Schielzeth, 2013). In addition, likelihood ratio tests (LRTs) were used to assess whether the inclusion of additional predictors significantly improved model fit. Lower AIC and BIC values, along with higher LL and R² values, indicate better model–data fit. Conclusions, Expected Outcomes or Findings Based on the model comparison results, Model 3 was identified as the model providing the best fit to the data. Likelihood ratio tests (LRTs) indicated that each model differed significantly from its preceding model in the hierarchical sequence. For Model 3, the explained variance was R² = 27.8%, indicating that the predictors included in the model accounted for 27.8% of the variance in students’ probability of correctly answering CTT items. These findings suggest that the probability of a correct response is jointly determined by both item-level characteristics and student-level characteristics. Item linguistic complexity substantially reduced the probability of a correct response (OR = 0.517, p < .001), indicating that the linguistic structure of the item text constitutes a barrier to student performance and represents a source of CIV. However, this effect varied significantly as a function of student characteristics. The Complexity × Reading Achievement interaction was negative (OR = 0.942, p < .001), suggesting that as reading proficiency increases, the detrimental effect of linguistic complexity decreases. Similarly, the negative effect of complexity was more pronounced for students with a different home language background (OR = 0.869, p < .001). The Complexity × Homework interaction was also negative (OR = 0.958, p < .001), indicating that linguistically complex items are relatively more challenging for students who spend more time on homework in the test language. In contrast, the Complexity × Gender interaction was positive (OR = 1.158, p < .001), suggesting that the negative impact of linguistic complexity is comparatively weaker for male students. The Complexity × ESCS interaction was not statistically significant. In summary, reading achievement and gender emerge as mitigating factors that attenuate the negative effect of linguistic complexity, whereas language background intensifies this effect. The absence of a significant interaction between linguistic complexity and socioeconomic status suggests that linguistic barriers function as a direct cognitive constraint, largely independent of social class differences. References Abedi, J. (2002). Standardized achievement tests and English language learners: Psychometrics issues. Educational Assessment, 8(3), 231–257. https://doi.org/10.1207/S15326977EA0803_02 Abedi, J. (2004). The No Child Left Behind Act and English language learners: Assessment and accountability issues. Educational Researcher, 33(1), 4–14. https://doi.org/10.3102/0013189X033001004 Abedi, J. (2006). Language issues in item development. In S. M. Downing & T. M. Haladyna (Eds), Handbook of Test Development (pp. 377–398). Lawrence Erlbaum Associates. Abedi, J., & Lord, C. (2001). The language factor in mathematics tests. Applied Measurement in Education, 14(3), 219–234. https://doi.org/10.1207/S15324818AME1403_2 Abedi, J., Lord, C., Hofstetter, C., & Baker, E. (2000). Impact of accommodation strategies on English language learners’ test performance. Educational Measurement: Issues and Practice, 19(3), 16–26. https://doi.org/10.1111/j.1745-3992.2000.tb00034.x Bulut, H. C., Bulut, O., & Arikan, S. (2023). Evaluating group differences in online reading comprehension: The impact of item properties. International Journal of Testing, 23(1), 10–33. https://doi.org/10.1080/15305058.2022.2044821 Messick, S. (1994). Foundations of validity: Meaning and consequences in psychological assessment. ETS Research Report Series, 2. https://doi.org/10.1002/j.2333-8504.1993.tb01562.x Nakagawa, S., & Schielzeth, H. (2013). A general and simple method for obtaining R2 from generalized linear mixed‐effects models. Methods in Ecology and Evolution, 4(2), 133-142. https://doi.org/10.1111/j.2041-210x.2012.00261.x Renaissance Learning. (2012). Text complexity: Accurate estimates and educational recommendations. Wisconsin Rapids. https://renaissance.widen.net/view/pdf/ob12ylc2hr/R54882.pdf?t.download=true&u=zceria Rubin, D. B. (1987). Multiple imputation for nonresponse in surveys (1. ed.). Wiley. https://doi.org/10.1002/9780470316696 | ||