Conference Agenda
| Session | ||
09 SES 14 B: Country-level Patterns Behind Mathematic Achievement Using TIMSS and PISA
Paper Session | ||
| Presentations | ||
09. Assessment, Evaluation, Testing and Measurement
Paper What Shapes Mathematics Literacy in Times of Crisis? Multilevel Evidence from Post-Pandemic International Assessment Data in Türkiye Gaziantep University, Turkey (Türkiye) Presenting Author:Large-scale assessments provide comprehensive data on the effectiveness of educational policies, allowing countries to evaluate the current state of their education systems. Fundamental elements such as the impact of instructional methods on students, learning barriers, and the contribution of curricula to student achievement are systematically examined through these measurement tools (Erkuş, 2012). In this context, the Programme for International Student Assessment (PISA) conducted by the Organisation for Economic Co-operation and Development (OECD), serves as a critical international comparison tool that enables countries to review their policies and create strategic roadmaps (Ministry of National Education [MoNE], 2019).PISA assesses not only whether 15-year-old students can reproduce knowledge but also how they extrapolate from what they have learned and apply their knowledge in unfamiliar settings, both in and out of school (OECD, 2023). Beyond describing achievement patterns, large-scale assessments such as PISA play a central role in monitoring educational quality and informing system-level improvement strategies, particularly in periods of systemic disruption such as the post-pandemic era. In the PISA 2022 cycle, mathematics literacy was the major domain. Mathematics literacy is defined as the capacity to formulate, employ, and interpret mathematics in a variety of contexts to describe, explain, and predict phenomena (MoNE, 2023). According to PISA 2022 results, Türkiye ranked 39th among 81 countries with a score of 453; while this indicates an improvement compared to 2018, it remains significantly below the OECD average of 472 (Aksakal, 2024). Social science research often involves investigating relationships between individuals and the environments in which they learn, live, and work (Hox et al., 2018). PISA data is inherently nested, with students clustered within schools. Consequently, traditional single-level analyses that ignore this hierarchical structure often disaggregate group-level variables to the individual level or aggregate individual-level variables to the group level (Hofmann, 1997). Such approaches violate independence assumptions and can lead to biased estimates (Julian, 2001; Muthén & Satorra, 1995). To address these methodological challenges, multilevel models provide a robust framework for simultaneously examining relationships at each level of the data hierarchy, thereby yielding more comprehensive and realistic results (Heck, 2001). This study aims to examine the structural relationships between students' mathematics literacy and student- and school-level factors in order to provide evidence relevant to educational improvement and quality assurance processes in the Turkish education system. Furthermore, given that PISA 2022 was administered in the aftermath of the COVID-19 pandemic, this research seeks to uncover the variables influencing the education system during this extraordinary period. Specifically, the study addresses the following research questions:
The conceptual framework of the study is informed by ecological and multilevel perspectives on learning. From an ecological standpoint, student achievement is shaped by multiple, nested environments, including classroom and school contexts. Methodologically, the study is grounded in multilevel modelling theory, which explicitly accounts for the nested structure of educational data (Hox et al., 2018). Multilevel SEM extends traditional hierarchical linear modelling by allowing the estimation of latent variables, measurement error, and complex structural relations across levels simultaneously (Heck, 2001; Heck & Thomas, 2020). These features are particularly important when analysing PISA data, which rely on plausible values derived from item response theory and require the appropriate application of sampling weights to ensure valid population-level inferences. Methodology, Methods, Research Instruments or Sources Used This study adopts a quantitative research design grounded in a relational survey model to examine the factors shaping mathematics literacy during crises. The analysis draws on secondary data from the PISA Türkiye dataset, which provides nationally representative information on students’ academic performance and contextual characteristics. Given the hierarchical structure of the PISA data—students nested within schools—and the presence of latent constructs measured with error, Multilevel Structural Equation Modeling (MSEM) is employed as the primary analytical approach. MSEM enables the simultaneous examination of relationships at both the student and school levels while explicitly accounting for measurement error in latent variables. Compared to conventional multilevel regression techniques, MSEM allows for the estimation of complex pathways, including direct and indirect effects among multiple variables, and offers greater flexibility in handling non-normal distributions and categorical indicators (Heck & Thomas, 2020). This makes MSEM particularly suitable for large-scale assessment data such as PISA. By linking multilevel modeling with latent variable analysis, the study provides an evidence-based framework that can support educational quality assurance and data-informed improvement strategies in post-crisis contexts. Sampling and Weighting Procedures PISA uses a two-stage stratified sampling design, in which schools are sampled first and students are selected within schools in the second stage. To ensure unbiased population estimates and to adjust for unequal probabilities of selection, both student-level and school-level sampling weights are incorporated into all analyses. This weighting procedure enhances the generalizability of the findings to the national student population. Measures Dependent Variable. Mathematics literacy serves as the primary outcome variable. In line with PISA technical guidelines, achievement is represented using ten plausible values rather than a single point estimate. This approach, grounded in Item Response Theory, captures measurement uncertainty and yields more reliable population-level inferences. Independent Variables. The explanatory variables are derived from student and school questionnaires and reflect key contextual dimensions relevant to educational outcomes in crisis contexts. These include indicators of socio-economic background, school climate, and instructional quality. Together, these variables enable an examination of how individual and institutional factors jointly contribute to mathematics literacy in the post-pandemic period. By integrating multilevel modeling with latent variable analysis, this methodological framework provides a robust basis for understanding the multiscalar determinants of mathematics literacy within the Turkish education system. Conclusions, Expected Outcomes or Findings This research is anticipated to be one of the pioneering studies employing Multilevel Structural Equation Modeling (MSEM) on the PISA 2022 Türkiye dataset. Existing literature on PISA data analysis in Türkiye has often neglected critical methodological components such as sampling weights and plausible values. By rigorously applying these elements within an MSEM framework, this study aims to set a methodological standard for future large-scale assessment research, ensuring high validity and reliability of the findings. Substantively, the study expects to reveal the multi-faceted impact of the COVID-19 pandemic on mathematics achievement. By modeling the relationships between student background, school climate, teacher support, and literacy scores, the findings will identify which factors acted as buffers or barriers during the pandemic. The decomposition of variance into student and school levels will provide policymakers with evidence-based insights on whether interventions should target school-wide systemic changes or individual student support mechanisms. Ultimately, the results will contribute to the development of resilient education policies capable of mitigating the effects of future crises. References Aksakal, D. (2024). PISA 2022 sonuçlarına göre Türkiye başarısının değerlendirilmesi [Evaluation of Türkiye's success according to PISA 2022 results] [Unpublished master's thesis]. Çanakkale Onsekiz Mart University. Erkuş, A. (2012). Psikolojide ölçme ve ölçek geliştirme [Measurement and scale development in psychology]. Pegem Akademi. Heck, R. H. (2001). Multilevel modeling with SEM. In New developments and techniques in structural equation modeling(pp. 109-148). Psychology Press. Heck, R., & Thomas, S. L. (2020). An introduction to multilevel modeling techniques: MLM and SEM approaches. Routledge. Hofmann, D. A. (1997). An overview of the logic and rationale of hierarchical linear models. Journal of management, 23(6), 723-744. Hox, J.J., Moerbeek, M & Van de Schoot, R. (2018). Multilevel analysis: Techniques and applications. Routledge. Julian, M. W. (2001). The consequences of ignoring multilevel data structures in nonhierarchical covariance modeling. Structural equation modeling, 8(3), 325-352. Ministry of National Education. (2023). PISA 2022 Türkiye report. Ministry of National Education. Ministry of National Education. (2019). PISA 2018 Türkiye ön raporu [PISA 2018 Türkiye preliminary report]. Ministry of National Education. Muthén, B. O., & Satorra, A. (1995). Complex sample data in structural equation modeling. Sociological methodology, 267-316. 09. Assessment, Evaluation, Testing and Measurement
Paper Patterns of Association with Grade 4 Mathematics Achievement in the UAE and GCC: Evidence from TIMSS 2023 Emirates College for Advanced Education, United Arab Emirates Presenting Author:Using Grade 4 data from the digitally administered TIMSS 2023 assessment, this study examines whether key instructional and motivational correlates of mathematics performance show similar patterns across education systems by comparing students in the United Arab Emirates (UAE) with students in four other GCC countries (Bahrain, Oman, Qatar, and Saudi Arabia). Comparing the UAE with other GCC countries is informative because these systems share a broadly common regional policy context yet differ in how schooling is organized and resourced. Participation in TIMSS supports cross-system comparability in both achievement and questionnaire-based constructs (Mullis et al., 2021; von Davier, 2024). GCC countries nevertheless vary in reform trajectories and modernization efforts, which may shape how classroom experiences translate into competence beliefs and performance (World Bank, 2024). They also differ in demographic composition, such as the size and diversity of migrant and expatriate populations, which can influence school contexts and how self-belief measures function across settings. In addition, variation in digitalization and technology integration in schooling provides an important contextual background for interpreting how digital competence beliefs relate to learning outcomes (OECD, 2023). Overall, a UAE–other GCC comparison is both theoretically meaningful and policy-relevant for examining when digital self-efficacy and instructional clarity are more strongly linked to mathematics confidence and achievement. The study focuses on three policy-relevant constructs available in TIMSS: instructional clarity in mathematics (students’ perceptions of clear and understandable instruction), digital self-efficacy (students’ perceived competence in using digital tools for learning), and confidence in mathematics (students’ competence beliefs). The conceptual framing relies primarily on Social Cognitive Theory (Bandura, 1997), which emphasizes self-efficacy beliefs as key determinants of learning-related behaviors and achievement, and Expectancy–Value Theory (Eccles & Wigfield, 2002), which suggests that competence beliefs (expectancies for success) shape engagement and performance outcomes. In addition, the focus on instructional clarity is consistent with research on effective teaching and teacher clarity as a predictor of student learning (Hattie, 2009). A multi-group mediation (path) model is used to test whether confidence in mathematics functions as a mechanism linking instructional clarity and digital self-efficacy to mathematics achievement, and whether the strength of these direct and indirect pathways differs across the UAE and GCC contexts. This mediation logic is aligned with contemporary methodological guidance that emphasizes estimating and testing indirect effects directly in path/SEM frameworks (MacKinnon, 2008; Hayes, 2018). Because cross-system comparisons can mask meaningful contextual differences, the study uses multi-group SEM to estimate group-specific structural relations and formally compare pathways across contexts (Byrne, 2016; Kline, 2016). Gender is included as a covariate predicting both confidence and achievement in each group to account for potential gender-related differences in competence beliefs and performance (Else-Quest et al., 2010). Four research questions guide the study: (1) How do digital self-efficacy and instructional clarity in mathematics predict Grade 4 mathematics achievement? (2) To what extent does confidence in mathematics mediate the relationships between digital self-efficacy and achievement, and between instructional clarity and achievement? (3) Do the direct structural relations among digital self-efficacy, instructional clarity, confidence, and achievement differ between the UAE and GCC groups? (4) Do the indirect effects via confidence differ between groups? Corresponding hypotheses specify positive direct effects of digital self-efficacy and instructional clarity on achievement and on confidence, a positive confidence-to-achievement link, and cross-group differences in selected pathways. This proposal leverages TIMSS 2023 to test whether instructional and motivational pathways to mathematics achievement generalize across education systems, and to identify where associations differ by context. Findings can inform the interpretation of cross-system comparisons and efforts to strengthen early mathematics learning. Methodology, Methods, Research Instruments or Sources Used Data come from the Trends in International Mathematics and Science Study (TIMSS) 2023 Grade 4 digitally administered mathematics assessment in the United Arab Emirates (UAE) and four other Gulf Cooperation Council (GCC) education systems (Bahrain, Oman, Qatar, and Saudi Arabia (IEA, 2025). The dataset included 60,272 students (UAE = 34,842; other GCC = 25,430); two cases were excluded from the SEM analyses because they had missing data on the exogenous predictor gender (ITSEX), yielding a final sample of 60,270. Measures Mathematics achievement was represented by TIMSS five plausible values (PVs), which provide multiple imputed draws of students’ latent mathematics proficiency for valid population inference under large-scale assessment designs (Mislevy et al., 1992). Analyses were conducted separately for each PV dataset, and estimates were pooled using Rubin’s rules to incorporate within- and between-imputation variability (Rubin, 1987). Non-achievement constructs, including digital self-efficacy, instructional clarity in mathematics, and confidence in mathematics, were TIMSS student questionnaire indices derived from multiple items using TIMSS scaling procedures and used as observed variables in the analysis (Fishbein et al., 2025). Analytic approach Because TIMSS uses a complex, multistage clustered sampling design with unequal selection probabilities, all analyses incorporated the TIMSS full-sample student weight and JK2 replicate weights to obtain design-consistent standard errors (IEA, 2025). Models were estimated in Mplus 8.8 using maximum likelihood with TYPE = COMPLEX and REPSE = JACKKNIFE2 (250 replicates; Muthén & Muthén, 1998–2017). Missing data on questionnaire variables were handled using full information maximum likelihood under a missing-at-random framework (Enders, 2010). A multi-group SEM compared the UAE with other GCC systems. Confidence in mathematics was specified as a mediator linking digital self-efficacy and instructional clarity to achievement, with digital self-efficacy and instructional clarity allowed to covary. Gender was included as a covariate predicting both confidence and achievement in each group. PV-specific results were pooled in R (v4.4.1; R Core Team, 2024) using Rubin-style variance rules proposed by Rubin (1987). Between-group differences in direct and indirect effects were tested using t-based inference after pooling, with Satterthwaite degrees of freedom approximations used for cross-group parameter comparisons (Satterthwaite, 1946). Conclusions, Expected Outcomes or Findings Across the UAE and other GCC systems, results show a clear instructional–motivational pattern. Mathematics confidence is a strong predictor of Grade 4 mathematics achievement. Instructional clarity in mathematics and digital self-efficacy are also positively related to achievement. These links occur both directly and indirectly through mathematics confidence. Descriptively, UAE students have a higher average mathematics achievement than the GCC group. UAE students also report slightly higher mathematics confidence, digital self-efficacy, and instructional clarity. Bivariate correlations are positive and statistically significant in both groups. Mathematics achievement is most strongly associated with mathematics confidence. The multi-group SEM converged across all plausible-value runs. Model misfit was low (SRMR = 0.022). Several effects differ by context. The digital self-efficacy → achievement path is stronger in the UAE than in the GCC. The predictor-to-confidence paths are also stronger in the UAE. Mediation is present in both groups. However, the indirect effects via confidence are larger in the UAE. Gender patterns are context-sensitive. Gender predicts higher confidence in the UAE but lower confidence in the GCC. This suggests that confidence formation differs across systems, even when the confidence–achievement link is similar. Overall, findings highlight mathematics confidence as a key pathway through which classroom experiences and students’ digital readiness relate to early mathematics achievement across systems. At the same time, the stronger UAE effects suggest that system context may shape how digital self-efficacy and instructional clarity translate into confidence and, in turn, achievement. References Bandura, A. (1997). Self-efficacy: The exercise of control. Freeman. Byrne, B. M. (2016). Structural equation modeling with AMOS: Basic concepts, applications, and programming (3rd ed.). Routledge. Eccles, J. S., & Wigfield, A. (2002). Motivational beliefs, values, and goals. Annual Review of Psychology, 53, 109–132. Else-Quest, N. M., Hyde, J. S., & Linn, M. C. (2010). Cross-national patterns of gender differences in mathematics: A meta-analysis. Psychological Bulletin, 136(1), 103–127. https://doi.org/10.1037/a0018053 Enders, C. K. (2010). Applied missing data analysis. Guilford Press. ttps://librarysearch.ed.ac.uk/discovery/fulldisplay?docid=alma992679763502466&context=L&vid=44UOE_INST:44UOE_VU2&lang=en Fishbein, B., Taneva, M., & Kowolik, K. (2025). TIMSS 2023 user guide for the international database. TIMSS & PIRLS International Study Center, Boston College; International Association for the Evaluation of Educational Achievement (IEA). https://timss2023.org/data Hattie, J. (2009). Visible learning: A synthesis of over 800 meta-analyses relating to achievement. Routledge. International Association for the Evaluation of Educational Achievement (IEA). (2025). TIMSS 2023 International Databas. https://www.iea.nl/international-data-tools-iea-studies/timss/timss-2023 Kline, R. B. (2016). Principles and practice of structural equation modeling (4th ed.). Guilford Press. Mislevy, R. J., Beaton, A. E., Kaplan, B., & Sheehan, K. M. (1992). Estimating population characteristics from sparse matrix samples of item responses. Journal of Educational Measurement, 29(2), 133–161. Mullis, I. V. S., Martin, M. O., & von Davier, M. (Eds.). (2021). TIMSS 2023 assessment frameworks. TIMSS & PIRLS International Study Center, Boston College; International Association for the Evaluation of Educational Achievement (IEA). Muthén, L. K., & Muthén, B. O. (1998–2017). Mplus user’s guide (8th ed.). Muthén & Muthén. https://www.statmodel.com/download/usersguide/MplusUserGuideVer_8.pdf Organisation for Economic Co-operation and Development. (2023). OECD Digital Education Outlook 2023: Towards an effective digital education ecosystem. OECD Publishing. https://doi.org/10.1787/c74f03de-en R Core Team. (2024). R: A language and environment for statistical computing. R Foundation for Statistical Computing. Vienna, Austria. https://www.r-project.org/ Satterthwaite, F. E. (1946). An approximate distribution of estimates of variance components. Biometrics Bulletin, 2(6), 110–114. von Davier, M., Fishbein, B., & Kennedy, A. (Eds.). (2024). TIMSS 2023 Technical Report (Methods and Procedures). TIMSS & PIRLS International Study Center, Boston College. https://timss2023.org/methods/ World Bank. (2024). Gulf Economic Update (Issue 12): Unlocking Prosperity: Transforming Education for Economic Breakthrough in the GCC. World Bank. 09. Assessment, Evaluation, Testing and Measurement
Paper Cracking the Code: Machine Learning Insights on School Factors Affecting PISA 2022 Math Scores in Türkiye Middle East Technical University, Turkey (Türkiye) Presenting Author:This proposal examines which school-level conditions are most strongly linked to mathematics success in Türkiye in PISA 2022. It uses the PISA 2022 School Questionnaire and supervised machine learning to model school success from leadership practices, school climate, resources, and family–school relations.
Research question Which school-level factors best predict whether a school in Türkiye is “successful” in mathematics in PISA 2022, and how well can a supervised model classify school success using school-level indicators?
Objective The objective is to build an interpretable predictive model and to identify actionable predictors that school leaders and policymakers can strengthen, while keeping sight of schools’ historical and social mission, as emphasized by Guevara-Reyes and his colleagues (2025).
Conceptual and theoretical framework · Leadership and administration. School leaders shape learning conditions by setting goals, managing resources, and building a climate conducive to learning. Leadership influences student outcomes directly and indirectly through teacher quality, instructional coherence, and equitable learning climates (Leithwood & Jantzi, 2006; Robinson et al., 2008). In mathematics, instructional leadership, distributed leadership, and data-driven decision making are positively related to student performance (Hallinger, 2011). In developing countries, leaders often work under limited resources, unclear policy mandates, and fragmented teacher training systems, so leadership capacity matters for linking policy and classroom practices to improve attainment (Bush, 2009). Administrators can also support teacher collaboration and community engagement, which is an indirect predictor of mathematics performance (OECD, 2023). · School climate. School climate (relationships, safety, and institutional values) shapes engagement and academic outcomes (Thapa et al., 2013). Self-Determination Theory and Stage-Environment Fit Theory emphasize how supportive contexts can meet students’ needs for competence, autonomy, and relatedness; positive teacher–student relationships can also buffer bullying and academic risks (Ren, Chen, & Zhao, 2025). Bullying undermines achievement by weakening students’ social bonds and sense of safety (Ren et al., 2025; Graham, 2016). International PISA evidence also shows that chronic absenteeism rose to 7.6% in PISA 2022, up from 4.7% in 2018 (OECD, 2023). · Resources and equity. Cross-country PISA research suggests that school-to-school achievement gaps are smaller where resources are distributed more equally (Beese & Liang, 2010). School material resources and the quality of the learning environment show meaningful links to achievement in PISA-based studies (Trinidad, 2020). School-level safety, instructional leadership, and parental support also predict mathematics performance (Wardat et al., 2022; Tomul, Önder, & Taşlıdere, 2021).
European/international dimension PISA is designed for international comparison, so identifying school-level predictors in Türkiye contributes to European and international discussions on equity, governance, and school improvement. The study defines school “Success” using an OECD-linked mathematics proficiency threshold used in OECD contexts (European Schools, 2022, p. 35), supporting interpretation across systems. The predictors prioritized by the model—parental involvement, disciplinary climate, technology use, leadership practices, student–teacher relationships, student absenteeism, and bullying—are widely relevant for equity-oriented school improvement efforts.
Method overview The study analyzes 81 school-level features from 196 Turkish schools in the PISA 2022 School Questionnaire. To address class imbalance and small sample size, it applies Synthetic Minority Over-sampling Technique (SMOTE) and creates 3,564 balanced observations (Chawla et al., 2002). It tests supervised classifiers and reports logistic regression results; the final sample meets recommended minimums for stable modeling (Silvey & Liu, 2024). Generalizability is assessed using the training–test accuracy gap criterion (Géron, 2022).
Expected contribution The study shows how interpretable machine learning can produce policy-relevant evidence from large international assessment data and highlight institutional levers linked to mathematics success (Acısıl-Çelik & YeşilKana, 2022; Erdoğan &Taştan, 2024). Methodology, Methods, Research Instruments or Sources Used This study employed a predictive modeling research design to identify school-level factors associated with mathematics achievement in Türkiye, analyzing publicly available OECD PISA 2022 School Questionnaire data. The Türkiye sample included 196 schools, each described by 81 school-level features. Due to the small sample size and class imbalance, the SMOTE was applied to augment and balance the data, resulting in 3,564 school-level observations (Chawla et al., 2002). Data preprocessing was used to improve data quality and suitability for modeling. Missing data were handled through imputation: discrete variables were imputed using the mean when values were close to the median, and the median when distributions were skewed. For ordinal Likert-type variables, the mode was used to preserve the original category order. Outliers in numeric variables were detected with the Interquartile Range (IQR) method. Except for the outcome variable (“Success”), outliers were capped through Winsorization to reduce skewness. After cleaning, box plots were generated for each numeric feature to review distributions and check whether outlier treatment worked as intended. The outcome variable “Success” was coded as 1 for scores above 65 and as 0 for scores below 65, in line with the 65% math proficiency threshold in OECD countries (European Schools, 2022, p. 35). After balancing with SMOTE, the dataset was randomly divided into training (80%) and test (20%) sets. For modeling, two supervised classification algorithms were tested: Random Forest (RF) and Logistic Regression (LR). Model performance was evaluated using accuracy and weighted F1-score. LR outperformed RF in this context, so only the LR results were reported. The final sample size (3,564 cases) exceeds recommended minimums for stable modeling (696 for LR and 3,404 for RF) (Silvey & Liu, 2024). Potential overfitting was evaluated by examining the accuracy gap between the training and test sets; a difference below 5% was treated as evidence of satisfactory generalizability (Géron, 2022). Conclusions, Expected Outcomes or Findings This study shows that school-level indicators from PISA 2022 can predict mathematics “Success” in Türkiye with high accuracy. After cleaning the data and balancing the classes with SMOTE (Chawla et al., 2002), the Logistic Regression (LR) model achieved training accuracy of 0.9161 and test accuracy of 0.9038, with weighted F1-scores of 0.9160 (training) and 0.939 (test). The gap between training and test accuracy was well below 5%, indicating strong generalizability and no evidence of overfitting (Géron, 2022). In terms of findings, parental involvement emerged as the strongest predictor of mathematics achievement, followed by disciplinary climate. Other important predictors included technology use, school background, student–teacher relationships, student absenteeism, administrative behavior, student support, bullying, leadership activities, and leadership practices. Overall, the results underline that mathematics success is linked to a combination of relational, organizational, and resource-related conditions inside schools. Expected outcomes include a clearer, evidence-based picture of which institutional levers matter most for school improvement efforts in Türkiye. By highlighting actionable predictors, the study offers a practical basis for school reform, leadership capacity building, and targeted interventions. More broadly, it demonstrates how machine learning can turn large international datasets into scalable, policy-relevant insights for educational decision-making (Acısıl-Çelik & YeşilKana, 2022; Erdoğan &Taştan, 2024). References Acıslı-Celik, S., & Yesilkanat, C.M. (2023). Predicting science achievement scores with machine learning algorithms: A case study of OECD PISA 2015–2018 data. Neural Computing and Applications,35(28),21201-21228. Beese, J.A., & Liang, X. (2010). Do resources matter? PISA science achievement comparisons between students in the United States, Canada and Finland. Improving Schools,13(3),266–279.https://doi.org/10.1177/1365480210378941 Bush, T. (2009). Leadership development and school improvement: Contemporary issues in leadership development. Educational Review,61(4),375–389.https://doi.org/10.1080/00131910903403956 Chawla, N.V., Bowyer, K.W., Hall, L.O., & Kegelmeyer, W.P. (2002). SMOTE: Synthetic Minority Over-sampling Technique. Journal of Artificial Intelligence Research,16,321-357.https://doi.org/10.1613/jair.953 Erdoğan, S., & Taştan, H. (2024). Predicting Student Achievement via Machine Learning: Evidence from Turkish Subset of PISA. Yildiz Social Science Review,10(1),7-27. European Schools. (2022). PISA-based Test for European Schools: Group Report 2022. https://www.eursc.eu/Documents/Group%20Report_PISA_based_Test_for_European_schools_2022_en.pdf Géron, A. (2022). Hands-on machine learning with scikit-learn, keras and tensorflow: Concepts, tools, and techniques to build Intelligent Systems. O’Reilly Media, Inc. Guevara-Reyes, R., Ortiz-Garcés, I., Andrade, R., Cox-Riquetti, F., & Villegas-Ch, W. (2025). Machine learning models for academic performance prediction: Interpretability and application in educational decision-making. Frontiers in Education,10.https://doi.org/10.3389/feduc.2025.1632315 Hallinger, P. (2011). Leadership for learning: Lessons from 40 years of empirical research. Journal of Educational Administration,49(2),125–142.https://doi.org/10.1108/09578231111116699 Leithwood, K., & Jantzi, D. (2006). Transformational School Leadership for large-scale reform: Effects on students, teachers, and their classroom practices. School Effectiveness and School Improvement,17(2),201–227.https://doi.org/10.1080/09243450600565829 OECD. (2023). PISA 2022 results: Learning during and after COVID-19. OECD Publishing. Ren, Y., Chen, X., & Zhao, J. (2025). The protective role of teacher-student relationships and school belonging against bullying in Chinese middle schools. Journal of School Psychology,93,112–124. Robinson, V.M., Lloyd, C.A., & Rowe, K.J. (2008). The impact of leadership on student outcomes: An analysis of the differential effects of leadership types. Educational Administration Quarterly,44(5),635–674.https://doi.org/10.1177/0013161x08321509 Silvey, S., & Liu, J. (2024). Sample size requirements for popular classification algorithms in tabular clinical data: Empirical study. Journal of Medical Internet Research,26.https://doi.org/10.2196/60231 Thapa, A., Cohen, J., Guffey, S., & Higgins-D’Alessandro, A. (2013). A review of school climate research. Review of Educational Research,83(3),357–385. Tomul, E., Önder, E., & Taşlıdere, E. (2021). The impact of school-related variables on mathematics achievement in Turkey: A regional comparison. Educational Studies,47(6),637–655. Trinidad, J. E. (2020). Material Resources, school climate, and achievement variations in the Philippines: Insights from Pisa 2018. International Journal of Educational Development,75,102174.https://doi.org/10.1016/j.ijedudev.2020.102174 Wardat, Y., Almekhlafi, A., & BinSubaih, A. (2022). School factors influencing students' mathematics performance in the UAE. International Journal of Educational Research,112,101912. 09. Assessment, Evaluation, Testing and Measurement
Paper Predicting Students’ Mathematics Literacy with Educational Data Mining: A Cross-National Analysis of PISA 2022 Gazi University, Turkey (Türkiye) Presenting Author:Problem Statement Large-scale assessments are standardized evaluations conducted at regional, national, or international levels with large student populations. They support teaching and learning by enabling cross-system comparisons and providing evidence for policy decisions such as curriculum reform, teacher professional development, and equitable resource allocation (Clarke & Luna-Bazaldua, 2021; Luo, 2023; Tan et al., 2024; Zhu et al., 2025). Among these assessments, the Programme for International Student Assessment (PISA), coordinated by the OECD, is one of the most influential. Conducted every three years since 2000, PISA measures the reading, mathematics, and science literacy of 15-year-olds worldwide (Luo, 2023; OECD, 2023a; Tan et al., 2024; Zhu et al., 2025). The PISA 2022 cycle focused on mathematics literacy, defined as the ability to reason mathematically and to apply mathematics to real-life contexts (OECD, 2023a). Persistent cross-national differences in achievement suggest that mathematics literacy is shaped by educational systems, cultural orientations, and inclusion policies (Bertling & Alegre, 2019; OECD, 2023a). Prior research has examined multiple predictors of achievement, including motivation-related factors (e.g., self-efficacy, self-concept, interest, anxiety), language background, health-related factors, socio-economic status, safety and bullying, family resources, and teacher-related factors such as guidance, fairness, feedback, and experience (Martínez-Abad & Chaparro-Caso-López, 2016; Tan et al., 2024; Wan et al., 2025; Wang et al., 2023; Zhu et al., 2025). However, multidimensional constructs such as immigration and language exposure, school culture and climate, ICT familiarity, assessment systems, teacher qualifications, and global crises remain underexplored despite their relevance for understanding equity and learning opportunities (Bertling & Alegre, 2019). This study focuses on Germany, Switzerland, and Türkiye due to their cultural and linguistic diversity in PISA 2022. Switzerland’s multilingual structure and substantial foreign-born population create unique educational dynamics (Eurydice, 2023), while Germany and Türkiye face comparable challenges related to immigration, language integration, and inclusion. These issues gained additional importance after the COVID-19 pandemic, which disrupted learning globally and highlighted inequalities in access, preparedness, and instructional continuity (Maldonado & De Witte, 2021; UNESCO, 2022). The pandemic also increased the relevance of teacher-developed assessment practices and teachers’ digital competence for sustaining mathematics learning (Scherer et al., 2021; Wellberg, 2023). Methodologically, PISA datasets are highly suitable for educational data-mining approaches because they include large samples, multiple countries, and a wide range of student, family, teacher, and school-level variables. Data mining, referred to as knowledge discovery in databases, involves identifying meaningful patterns, structures, and relationships within large and complex datasets (Kretowski, 2019). The purpose of this study is to compare the predictive performance of selected educational data-mining algorithms in classifying students’ PISA 2022 mathematics literacy performance and to identify the most influential predictors across Germany, Switzerland, and Türkiye. Also, this study aimed to examine whether PISA 2022 background variables (immigration and language exposure, school culture and climate, beliefs and attitudes, ESCS, ICT familiarity, teacher qualifications and professional development, assessment practices, and global crises) contribute to predicting students’ mathematics literacy outcomes across Switzerland, Germany, and Türkiye. Accordingly, the study addressed (1) the classification performance of C5.0 decision tree, random forest, Naive Bayes, and artificial neural network algorithms in predicting students’ PISA 2022 mathematics literacy performance and (2) the most influential predictors across the three contexts. Methodology, Methods, Research Instruments or Sources Used Method This study employed a predictive relational research design to examine factors associated with students’ mathematics literacy achievement in PISA 2022 (Creswell, 2012). Using educational data mining techniques, data from 15.524 students in Switzerland, Germany, and Türkiye were analyzed in IBM SPSS Modeler 18.0. These three countries were selected to represent high (Switzerland), middle (Germany), and low (Türkiye) performance groups. Population and Sample The sample reflected clear contextual differences across countries. Switzerland displayed the highest linguistic diversity, as students completed the assessment in German (f = 2.674; 59.8%), French (f = 1.197; 26.7%), or Italian (f = 604; 13.5%), whereas Germany (German: f = 3.964; 100%) and Türkiye (Turkish: f = 7,085; 100%) administered the assessment in a single language. The proportion of native students was lowest in Switzerland (f = 2,963; 66.2%) compared to Germany (f = 3.026; 76.3%) and Türkiye (f = 6.975; 98.4%). Home and test language differences were most common in Switzerland (f = 1.221; 27.3%), followed by Germany (f = 706; 17.8%) and Türkiye (f = 576; 8.1%). Data Collection and Variables The study used publicly available PISA 2022 mathematics literacy scores and background questionnaire scales from the official OECD database. Mathematics literacy achievement was treated as the dependent categorical variable. Independent variables included both categorical variables (LANG_DIFF, IMMIG) and continuous index variables computed by the PISA International Working Group. Data Analysis Missing data were examined prior to modeling. After excluding unanswered scales and consecutive missing responses, 15.254 observations and 28 variables were retained. The overall missing rate was 2.38%, and Little’s MCAR test indicated that data were missing completely at random (Wang et al., 2023). Random forest imputation was applied using R. Predictors were derived from weighted likelihood estimates (WLEs) and standardized for cross-country comparability. Five imputed values were generated, and their mean was used in the final dataset. The completed dataset was divided into training and test sets using a 70:30 split. Four algorithms were compared: C5.0 decision tree, random forest, Naive Bayes, and artificial neural network. Analyses were conducted separately for Türkiye, Germany, and Switzerland. Model performance was evaluated using classification metrics including accuracy, Cohen’s Kappa, AUC (Area Under the Receiver Operating Characteristic Curve), and MCC (Matthews Correlation Coefficient) (Bramer, 2020; Kretowski, 2019). Conclusions, Expected Outcomes or Findings The findings showed that the C5.0 algorithm produced the highest test-data classification performance in Türkiye and Germany. In Türkiye, C5.0 correctly classified 82.78% of students, with a Kappa value of 0.568, AUC of 0.864, and MCC of 0.570. In Germany, C5.0 achieved 75.46% accuracy, with a Kappa value of 0.505, AUC of 0.833, and MCC of 0.506. In Switzerland, Naive Bayes performed best, with 74.14% accuracy, Kappa of 0.478, AUC of 0.815, and MCC of 0.477. Random forest showed relatively high training performance but lower test performance, suggesting possible overfitting. In terms of predictor importance, mathematics self-efficacy emerged as the strongest predictor in all three countries. Teacher behaviors affecting school climate, teacher-developed tests, and the difference between home language and test language were also common predictors. Country-specific patterns were also observed. In Türkiye, pandemic-related school readiness, remote instruction capacity, digital preparedness, and mathematics teacher training were prominent. In Germany, socioeconomic status, language difference, immigration status, and multicultural views were more visible. In Switzerland, mathematics self-efficacy, socioeconomic status, teacher behavior, ICT feedback, digital preparedness, and language difference formed a more distributed predictor structure. References Bertling, J. & Alegre, J. (2019). PISA 2021 context questionnaire framework (field trial version). OECD Publishing. Bramer, M. (2020). Principles of data mining. London: Springer. Clarke, M., & Luna-Bazaldua, D. (2021). Primer on large-scale assessments o educational achievement. Washington: The World Bank. Crato, N., Patrinos, H.A. (2025). PIRLS 2021 and PISA 2022 Statistics Show How Serious the Pandemic Losses Are. In: Crato, N., Patrinos, H.A. (eds) Improving National Education Systems After COVID-19. Evaluating Education: Normative Systems and Institutional Practices. Springer, Cham. https://doi.org/10.1007/978-3-031-69284-0_1 Creswell, J. W. (2012). Educational research: Planning, conducting, and evaluating quantitative and qualitative research. Boston: Pearson. Eurydice (2023). Political, social and economic background and trends. Kretowski, M. (2019). Evolutionary decision trees in large-scale data mining. Studies in Big Data. Switzerland: Springer. Luo, S. (2023). Factors affecting English reading in Macao, Hong Kong, and Singapore: combining machine learning methods and hierarchical linear regressions using pisa 2018 data (Ph. D. Dissertation). University of Macau, China Maldonado, J. E., & De Witte, K. (2021). The effect of school closures on standardised student test outcomes. British Educational Research Journal. doi:10.1002/berj.3754 Martínez-Abad, F., & Chaparro-Caso-López, A. A. (2016). Data-mining techniques i detecting factors linked to academic achievement. School Effectiveness and School Improvement, 28(1), 39-55. Organisation for Economic Co-operation and Development. (2023a). PISA 2022 results (Volume I): The state of learning and equity in education. Paris: OECD Publishing. doi:10.1787/53f23881-en. Scherer, R., Howard, S. K., Tondeur, J., & Siddiq, F. (2021). Profiling teachers’ readiness for online teaching and learning in higher education: Who’s ready? Computers in Human Behavior, 118, 1-16. Tan, L., Chen, F., & Wei, B. (2024). Examining key capitals contributing to students’ science-related career expectations and their relationship patterns: A machine learning approach. Journal of Research in Science Teaching, 61(8), 1975–2010. https://doi.org/10.1002/tea.21939 UNESCO (2022). COVID-19 School health and safety protocols: Good practices and lessons learnt to respond to Omicron. https://unesdoc.unesco.org/ark:/48223/pf0000380400.locale=en Wang, F., King, R. B., & Leung, S. O. (2023). Why do East Asian students do so well in mathematics? A machine learning study. International Journal of Science and Mathematics Education, 21(3), 691–711. https://doi.org/10.1007/s10763-022-10262-w Wellberg, S. (2023). Teacher-made tests: Why they matter and a framework for analyzing mathematics exams. Assessment in Education: Principles, Policy & Practice, 1-23. doi:10.1080/0969594X.2023.2189565 Zhu, L., You, H., Hong, M., Fang, Z. (2025). Predictive insights into U.S. students’ mathematics performance on PISA 2022 using ensemble tree-based machine learning models. International Journal of Educational Research, 130, 1-15. https://doi.org/10.1016/j.ijer.2025.102537 | ||