Conference Agenda
| Session | ||
09 SES 10 C: Modelling Student Growth and Assessment Processes in School Contexts
Paper Session | ||
| Presentations | ||
09. Assessment, Evaluation, Testing and Measurement
Paper Assessing School Influence on Language Development: Comparing Student Selection Test and External Summative Assessment Data Nazarbayev Intellectual schools, Kazakhstan Presenting Author:In many education systems, schools are judged primarily by final exam results, with high scores often taken as indicators of quality. Yet this view is limited: some students enter with a strong educational background, while others begin at a lower level. Without distinguishing prior knowledge from progress made during schooling, outcomes may be misinterpreted and the actual contribution of schools overlooked. The theoretical framework draws on value‑added models of school effectiveness (Hanushek & Rivkin, 2010; Koedel & Betts, 2011), which emphasize the importance of separating students’ initial achievement from the contribution of schools to subsequent progress. This approach is widely used in European educational research to promote fairer evaluation of school performance. Rather than focusing only on raw results, it highlights growth indicators, as reflected in policies in the UK, the Netherlands, Belgium, and Ireland (Timmermans & Thomas, 2015; Leckie & Goldstein, 2019; Brown, McNamara & O’Hara, 2016). Research on language acquisition in school contexts (Baker, 2011) further underscores the role of instructional environments and pedagogical practices in shaping outcomes. This is especially relevant in multilingual education systems, where language development depends not only on individual ability but also on teaching quality and the broader learning context. The aim of this study is to evaluate the contribution of Nazarbayev Intellectual schools (NIS) to students’ development in Kazakh and Russian by comparing entry scores obtained in the Grade 7 selection test with exit scores achieved in the Grade 12 external summative assessment. NIS is the flagship of Kazakhstani education. Admission is highly competitive, with on average 10 applicants per place, and only the strongest students succeed (Nazarbayev Intellectual schools, 2024). The schools admit talented, motivated students with the goal of fostering their growth into critical thinkers and independent learners. NIS campuses are located across Kazakhstan (Toybazarova & Nazarova, 2018). This study addresses two research questions:
The first question examines students’ starting point, highlighting how selection test results capture prior knowledge rather than the effect of schooling. The second investigates how Grade 12 external assessments, aligned with international benchmarks, reflect cumulative progress and the added value of schools by distinguishing between initial ability and growth achieved through instruction. Together, these questions allow for a more accurate evaluation of school effectiveness in language development. By comparing entry and exit scores, the study shows why it is important to look beyond raw results. It also connects to wider European discussions about fair assessment, multilingual education, and how schools can be judged not only by final exams but by the progress they help students make. Methodology, Methods, Research Instruments or Sources Used This study uses a quantitative design to explore how NIS influences students’ language development in Kazakh and Russian. The analysis focuses on one cohort across twenty schools, including 1,754 students studying in Kazakh classes and 889 students in Russian classes. Two main data sources are used. The first is the Selection Test (2019) in Kazakh and Russian, taken at Grade 7 entry, which provides entry scores. The second is the External Summative Assessment (2025), conducted at the end of Grade 12 by an independent organization and aligned with international standards such as Cambridge International A Level, which provides exit scores in both Kazakh and Russian. To examine the relationship between entry and exit scores, the study applies three approaches: 1. Correlation analysis to explore the relationship between entry and exit scores. 2. Comparative analysis to highlight differences between Kazakh and Russian classes, showing how the language of instruction shapes learning trajectories. 3. Value added analysis to estimate the schools’ contribution by comparing entry and exit scores, distinguishing prior knowledge from progress made during schooling. Together, these methods provide a comprehensive picture of how NIS contributes to students’ language development. By integrating correlation, comparative, and value‑added perspectives, the study not only examines the link between entry and exit scores but also reveals important differences across Kazakh and Russian classes. Conclusions, Expected Outcomes or Findings The findings reveal a clear distinction between entry and exit scores in NIS. Students often enter with high scores, reflecting strong preparation and baseline proficiency. Yet by Grade 12, external assessment results are lower, suggesting that initial success does not guarantee equally strong graduation performance. Differences between Kazakh and Russian classes further illustrate how instructional language and school environment shape learning. Students in Kazakh classes show relatively stable progress with smaller gaps between entry and exit scores, while students in Russian classes display weaker outcomes and larger discrepancies. At entry, motivation is extremely high, which explains strong performance in Grade 7. By Grade 12, however, scores appear weaker in percentage terms. This gap may be explained by the different purposes of assessment, the intensive preparation effect for admission, and changes in motivation and workload over time. The weak correlation between entry and exit scores suggests that selection tests have limited predictive power for long term academic success. Overall, the study argues that schools should be evaluated not only by raw achievement scores but also by the added value they provide in fostering student growth. The observed differences are not signs of weakness but rather reflect the distinct roles of entry and exit assessments: one measures readiness, the other captures comprehensive development within a demanding, trilingual, internationally benchmarked system. References 1.Aknouch, L. (2025). Understanding Summative and Formative Assessment Dynamics. TESOL and Technology Studies, 6(2), 26-37. 2.Bairbayeva, S. (2023, October). Advancing Equity and Diversity: A Comprehensive Case Study of Nazarbayev Intellectual Schools in Secondary Education. In Publisher. agency: Proceedings of the 4th International Scientific Conference «Reviews of Modern Science»(October 19-20, 2023). Zürich, Switzerland, 2023. 256p (p. 13). Universität Luzern. 3.Baker, C. (2011). Foundations of bilingual education and bilingualism. Multilingual matters. 4.Brown, M., McNamara, G., & O’Hara, J. (2016). Quality and the rise of value-added in education: The case of Ireland. Policy Futures in Education, 14(6), 810-829. 5.Center for Pedagogical Measurements (2015). Model of External Summative Assessment: Methodological Guidelines. Retrieved from https://cpi-nis.kz/upload/userfiles/files/%D0%9C%D0%BE%D0%B4%D0%B5%D0%BB%D1%8C%20C%D0%9E%20%20%D0%BE%D1%82%2027_08_2015%20%E2%84%9643(1).pdf 6.Hanushek, E. A., & Rivkin, S. G. (2010). Generalizations about using value-added measures of teacher quality. American economic review, 100(2), 267-271. 7.Koedel, C., & Betts, J. R. (2011). Does student sorting invalidate value-added models of teacher effectiveness? An extended analysis of the Rothstein critique. Education Finance and policy, 6(1), 18-42. 8.Leckie, G., & Goldstein, H. (2019). The importance of adjusting for pupil background in school value‐added models: A study of Progress 8 and school accountability in England. British Educational Research Journal, 45(3), 518-537. 9.Leckie, G., & Prior, L. (2022). A comparison of value-added models for school accountability. School Effectiveness and School Improvement, 33(3), 431-455. 10.Levy, J., Brunner, M., Keller, U., & Fischbach, A. (2019). Methodological issues in value-added modeling: an international review from 26 countries. Educational Assessment, Evaluation and Accountability, 31(3), 257-287. 11.Nazarbayev Intellectual schools (2024). Annual report. Retrieved from https://www.nis.edu.kz/storage/files/01JVKVXA97T8Q3TRKPFKWS82C3.pdf 12.Nazarbayev Intellectual schools (2025). Rules of Competitive Selection. Retrieved from https://www.nis.edu.kz/en/page/konkurstyq-irikteu-erezeleri 13.OECD (2013). Synergies for Better Learning: An International Perspective on Evaluation and Assessment. Retrieved from https://www.oecd.org/en/publications/synergies-for-better-learning-an-international-perspective-on-evaluation-and-assessment_9789264190658-en.html 14.Timmermans, A. C., & Thomas, S. M. (2015). The impact of student composition on schools’ value-added performance: A comparison of seven empirical studies. School Effectiveness and School Improvement, 26(3), 487-498. 15.Toybazarova, N. A., & Nazarova, G. (2018). The modernization of education in Kazakhstan: trends, perspective and problems. Bulletin of National academy of sciences of the Republic of Kazakhstan, 6(376), 104-114. 09. Assessment, Evaluation, Testing and Measurement
Paper Socioeconomic Differences in the Development of Inductive Reasoning from Grades 3 to 7: A Longitudinal Study in Hungary 1: Institute of Education, University of Szeged; 2: Doctoral School of Education, University of Szeged; 3: MTA–SZTE Digital Learning Technologies Research Group, University of Szeged; 4: Budapest Institute for Policy Analysis Presenting Author:A large body of international research has consistently documented positive associations between family socioeconomic background and academic achievement. Meta-analytic evidence generally indicates small to moderate effect sizes, suggesting that SES is a meaningful, though not deterministic, contributor to students’ learning outcomes (Harwell et al., 2017; Liu et al., 2022; Korous et al., 2022). At the same time, substantial variation exists in the reported magnitude of SES effects. While some large-scale syntheses report correlations in the moderate range (r ≈ .22–.28; Liu et al., 2022), others identify considerably weaker associations, particularly when controlling for various individual and school-level factors (Kim et al., 2019; Song & Tan, 2025). These findings suggest that SES effects are shaped not only by family-level resources but also by broader institutional, cultural, and policy environments. Methodology, Methods, Research Instruments or Sources Used Sample The analyses were based on data from the Hungarian Educational Longitudinal Program, a nationwide research designed to monitor students’ learning trajectories and development from early primary school through lower secondary education (Csapó, 2014). The program follows students across Grades 1 to 8 and covers a broad range of cognitive and non-cognitive domains. The present study draws on three waves of data collection conducted when students were in Grades 3, 5, and 7. The final sample comprised 6,232 students (50.1% girls), attending 155 schools and 287 classes. The sampling design aimed to approximate representativeness at the regional, county, and settlement-type levels. Approximately 5% of the national student cohort at the target grade level was included. Schools served as the primary sampling units, and all participating institutions were mainstream primary schools. Assessment of inductive reasoning Students’ inductive reasoning was assessed using online tests consisting of figural and numerical series and analogy tasks. Across measurement occasions, different but anchor-linked test forms were employed to enable developmental comparisons over time. Student performance was placed on a common latent ability scale using Rasch model. The resulting scale demonstrated excellent reliability (EAP/PV=.94). The assessments were administered in the ICT labs of the participating schools using the eDia online assessment platform (Csapó & Molnár, 2019). The system relies on individual measurement identifiers to ensure student anonymity, and no personally identifiable information was collected during data collection. Handling of missing data To address missing responses resulting from the longitudinal design, multiple imputation was applied (Wijesuriya et al., 2025). A random forest–based chained equations procedure was used to generate 100 imputed datasets, allowing uncertainty associated with missing values to be appropriately reflected in subsequent analyses. Socioeconomic background measure Students’ socioeconomic background was operationalized using a composite latent construct reflecting multiple dimensions commonly emphasized in the literature. To achieve this, a two-level latent confirmatory factor model was estimated separately for each imputed dataset. The socioeconomic background factor captured shared variance across four domains: parental education, geographical disadvantage, material resources, and parental involvement. Parental education was indicated by the highest qualification levels of both parents. Geographical disadvantage was represented by a set of settlement-level socioeconomic indicators. Material resources were measured through a comprehensive inventory of household assets, while parental involvement was assessed using items capturing academically relevant support behaviours at home. Conclusions, Expected Outcomes or Findings In line with previous cross-sectional results (Molnár, 2025), we found a consistent and substantial association between students’ socioeconomic background and their level of inductive reasoning across all grade levels. Socioeconomic status showed moderate correlations with inductive reasoning performance (r = 0.32, 0.34, and 0.37, respectively), which are slightly higher than the effect sizes typically reported in recent meta-analyses (Liu et al., 2022; Korous et al., 2022). This pattern aligns with prior evidence and reflects the strong role of socioeconomic background in educational outcomes observed in Hungary (OECD, 2023). Across the full sample, students’ inductive reasoning improved by approximately half a standard deviation between measurement waves. Quartile-based analyses revealed pronounced and persistent socioeconomic differences at each grade level. In particular, the performance gap between students in the lowest and highest SES quartiles approached one standard deviation, corresponding to several years of developmental difference in inductive reasoning within the same grade level. Latent growth curve models further clarified these patterns. Socioeconomic background showed a strong association with the latent intercept (b=32.17, SE=1.20, p<0.01), indicating marked SES-related differences in initial levels of inductive reasoning by Grade 3. Although a small positive association between continuous SES and growth was observed (b=1.40, SE=0.47, p<0.01), quartile-based comparisons revealed no systematic differences in growth rates between socioeconomic groups. Thus, early socioeconomic differences remained stable over time, a finding consistent with prior research suggesting limited SES effects on developmental slopes (Molnár, 2025; Liu et al., 2022). Future research should examine whether these patterns hold when additional individual- and school-level factors are considered and when longer developmental spans are analyzed. The findings highlight the importance of early interventions aimed at reducing SES-related gaps in fundamental cognitive skills. Such efforts may contribute to improving educational equity not only in Hungary, but also in education systems facing similarly strong socioeconomic gradients. References Csapó, B. (2014). A szegedi iskolai longitudinális program [The Hungarian educational longitudinal program]. In J. Pál & Z. Vajda (Eds.), Szegedi Egyetemi Tudástár 7: Bölcsészet- és társadalomtudományok (pp. 117–166). Szegedi Egyetemi Kiadó. Csapó, B., & Molnár, G. (2019). Online diagnostic assessment in support of personalized teaching and learning: The eDia system. Frontiers in Psychology, 10, Article 1522. https://doi.org/10.3389/fpsyg.2019.01522 Hanushek, E. A., Woessmann, L. (2008). The role of cognitive skills in economic development. Journal of Economic Literature 46(3), 607–68. https://doi.org/10.1257/jel.46.3.607 Harwell, M., Maeda, Y., Bishop, K., & Xie, A. (2017). The Surprisingly Modest Relationship Between SES and Educational Achievement. The Journal of Experimental Education, 85(2), 197–214. https://doi.org/10.1080/00220973.2015.1123668 Kim, S. W., Cho, H., & Kim, L. Y. (2019). Socioeconomic status and academic outcomes in developing countries: A meta-analysis. Review of educational research, 89(6), 875-916. https://doi.org/10.3102/0034654319877155 Klauer, K. J., & Phye, G. D. (2008). Inductive reasoning: A training approach. Review of Educational Research, 78(1), 85–123. https://doi.org/10.3102/0034654307313402 Korous, K. M., Causadias, J. M., Bradley, R. H., Luthar, S. S., & Levy, R. (2022). A systematic overview of meta-analyses on socioeconomic status, cognitive ability, and achievement: The need to focus on specific pathways. Psychological reports, 125(1), 55-97. https://doi.org/10.1177/0033294120984127 Liu, J., Peng, P., Zhao, B., & Luo, L. (2022). Socioeconomic status and academic achievement in primary and secondary education: A meta-analytic review. Educational Psychology Review, 34(4), 2867-2896. https://doi.org/10.1007/s10648-022-09689-y Molnár, G. (2025). Development of 6-to 17-year-old students’ inductive reasoning: What is the most sensitive period?. Thinking Skills and Creativity, 56, 101745. https://doi.org/10.1016/j.tsc.2024.101745 OECD (2023). PISA 2022 Results (Volume I): The State of Learning and Equity in Education. OECD Publishing. https://doi.org/10.1787/53f23881-en Perret, P. (2015). Children's inductive reasoning: Developmental and educational perspectives. Journal of Cognitive Education and Psychology, 14(3), 389–408. https://doi.org/10.1891/1945-8959.14.3.389 Song, Q., & Tan, C. Y. (2025). The association between family socioeconomic status and academic achievement: new estimates using three-level meta-analysis of PISA 2009–2018 data. Comparative Education, 1–20. https://doi.org/10.1080/03050068.2025.2545659 Vo, D. V., & Csapó, B. (2022). Measuring inductive reasoning in school contexts: a review of instruments and predictors. International Journal of Innovation and Learning, 31(4), 506-525. https://doi.org/10.1504/IJIL.2022.123179 Wijesuriya, R., Moreno-Betancur, M., Carlin, J. B., White, I. R., Quartagno, M., & Lee, K. J. (2025). Multiple imputation for longitudinal data: A tutorial. Statistics in Medicine, 44(3–4), e10274. https://doi.org/10.1002/sim.10274 09. Assessment, Evaluation, Testing and Measurement
Paper Exploring Students’ Response Strategies on TIMSS 2023 Mathematics Items Using a Response Time-Based Mixture Item Response Theory Model 1: Aydın Adnan Menderes University, Turkey (Türkiye); 2: Canakkale Onsekiz Mart University, Turkey (Türkiye) Presenting Author:In the TIMSS (the Trends in International Mathematics and Science Study) 2023 computer-based assessment, student process data was provided along with other achievement and context questionnaires data. Process data (e.g., response time) provide empirical evidence about the cognitive states and behaviors that occur while working on a test item and that influence how the measured construct(s) translate into item performance (Goldhammer & Zehner, 2017). By revealing underlying response mechanisms beyond the test score, these processes provide evidence supporting the validity of intended interpretations of test scores. This source of validity evidence is emphasized in the Standards for Educational and Psychological Testing (American Educational Research Association, American Psychological Association, and National Council on Measurement in Education, 2014). Examining response processes in TIMSS, as a large-scale assessment, provides valuable insights into the cognitive processes underlying item responses and contributes to a better understanding of the meaning and interpretation of test scores. There are various approaches to examining response processes, but mixture item response theory (IRT) models have been widely used due to their strengths. Mixture IRT models combine features of traditional IRT models with latent class analysis (LCA) and assume the presence of hidden subpopulations in the data (Sen & Cohen, 2019), which makes them especially useful for examining students’ distinct response strategies. Huang, Luo, and Jeon (2025) reviewed previous models within the mixture IRT framework for studying item-response strategies and proposed a new response time-based mixture IRT model. They grouped existing approaches into two categories: examinee-level mixture models (e.g., Jeon & De Boeck, 2019) and examinee-by-item-level mixture models (e.g., Nagy & Robitzsch, 2021). While examinee-level models assume that test-takers employ the same response strategy throughout the assessment, examinee-by-item-level models allow test-takers to switch between response strategies (Huang et al., 2025). The examinee-by-item-level mixture model proposed by Huang et al. (2025) assumes the presence of multiple latent classes, including one regular ability-based class (modeled using a two-parameter logistic [2PL] IRT model) and one or more non-ability-based secondary classes (e.g., rapid guessing, knowledge retrieval), and incorporates response times to predict examinee-by-item class membership. This modeling approach allows examinees to switch between response strategies across items, which is likely in practice, as they may use different strategies depending on item properties (e.g., difficulty, complexity) or other factors (e.g., fatigue, lack of motivation). In this study, we aim to explore which response strategies eighth-grade students use when responding to TIMSS 2023 mathematics items in Türkiye and Finland, two countries with similar average mathematics achievement. We apply a response time-based mixture IRT model proposed by Huang et al. (2025) to both easy and difficult item groups to examine potential differences in response processes across item difficulty levels. Methodology, Methods, Research Instruments or Sources Used The study uses a subset of Finland and Türkiye TIMSS 2023 eighth-grade mathematics achievement and process data, which are publicly available from the International Association for the Evaluation of Educational Achievement (IEA) (https://www.iea.nl/data-tools/repository/timss). In TIMSS 2023, each student is randomly assigned to one of 14 test booklets, each consisting of two mathematics and two science item blocks. Each item block appears in two booklets and is paired with different item blocks (Mullis, Martin, & von Davier, 2021). Because each student receives a different set of items, we focus on item blocks to maintain a consistent item sequence and enable clearer interpretation of response processes. The assessment booklets are divided into two levels of difficulty: less difficult and more difficult (Mullis et al., 2021). We select two mathematics item blocks: one from the less difficult set (ME1) and one from the more difficult set (MD3). The ME1 block appears in Booklets 8 and 9, whereas the MD3 block appears in Booklets 2 and 5. The ME1 item block contains 14 items (scored 0/1), while the MD3 block contains 15 items (scored 0/1). TIMSS 2023 provides student process data files that include three types of variables related to students’ on-screen navigation during the assessment: total time on screen, time on first screen visit, and number of screen visits (Fishbein, Taneva, & Kowolik, 2025). To meet the requirements of response time-based modeling, we use total time on screen as the response time variable. For the analysis, we will estimate two models separately for each item block and each country: (1) a 2PL model (DeMars, 2010) as baseline model and (2) a two-class examinee-by-item mixture IRT model proposed by Huang et al. (2025). The mixture IRT model will include two latent classes: an ability-based class and a secondary (i.e., non-ability-based) response class. The latent class variable will be regressed on double-centered, log-transformed response times as suggested. All models will be estimated using maximum likelihood with robust standard errors (MLR). Analyses will be conducted using Mplus Version 9 (Muthén & Muthén, 1998–2017). Conclusions, Expected Outcomes or Findings We expect that the two-class examinee-by-item mixture IRT model will fit the data better than the standard two-parameter logistic (2PL) IRT model, suggesting the presence of a secondary response process in addition to the regular ability-based one. Response time deviations are anticipated to be associated with these response processes. In particular, shorter response times accompanied by a high probability of correct answers may reflect rapid knowledge retrieval strategies, whereas longer response times with a high probability of correct answers may reflect more effortful knowledge retrieval strategies (Huang et al., 2025). Through cross-country and cross-item-block comparisons, the results are expected to offer deeper insights into the underlying cognitive processes employed by students in the TIMSS 2023 mathematics assessment. Overall, the findings are expected to contribute to understanding how students engage with assessment tasks of varying difficulty and how response processes relate to score interpretation. References American Educational Research Association, American Psychological Association, & National Council on Measurement in Education (Eds.) (2014). Standards for Educational and Psychological Testing. American Educational Research Association. DeMars, C. (2010). Item response theory: Understanding statistics measurement. Oxford University. Fishbein, B., Taneva, M., & Kowolik, K. (2025). TIMSS 2023 User Guide for the International Database. Boston College, TIMSS & PIRLS International Study Center. https://timss2023.org/data Goldhammer, F., & Zehner, F. (2017). What to make of and how to interpret process data. Measurement: Interdisciplinary Research and Perspectives, 15(3–4), 128–132. https://doi.org/10.1080/15366367.2017.1411651 Huang, S., Luo, J. & Jeon, M. (2025). A response time-based mixture item response theory model for dynamic item-response strategies. Behavior Research Methods, 57(54). https://doi.org/10.3758/s13428-024-02555-5 Jeon, M., & De Boeck, P. (2019). An analysis of an item-response strategy based on knowledge retrieval. Behavior Research Methods, 51, 697–719 https://doi.org/10.3758/s13428-018-1064-1 Mullis, I. V., Martin, M. O., & von Davier, M. (2021). TIMSS 2023 Assessment Frameworks. International Association for the Evaluation of Educational Achievement (IEA). Muthén, L. K., & Muthén, B. O. (1998-2017). Mplus user’s guide (8th Edition). Muthén & Muthén. Nagy, G., & Robitzsch, A. (2021). A continuous HYBRID IRT model for modeling changes in guessing behavior in proficiency tests. Psychological Test and Assessment Modeling, 63(3), 361-395. Sen, S., & Cohen, A. S. (2019). Applications of Mixture IRT Models: A Literature Review. Measurement: Interdisciplinary Research and Perspectives, 17(4), 177–191. https://doi.org/10.1080/15366367.2019.1583506 09. Assessment, Evaluation, Testing and Measurement
Paper Understanding the Drivers of Student Literacy Growth in Kazakhstan: A Multilevel Value-Added Study 1: BI Education, Kazakhstan; 2: Y.Altynsarin National Academy of Education, Kazakhstan Presenting Author:This paper presents a large-scale institutional research project examining the drivers of student literacy growth in Kazakhstan. The study focuses on students in Grades 3, 4, and 7 across 13 schools within the BI Education network and aims to identify student-, teacher-, and family-level factors associated with literacy development over an academic year. While literacy outcomes are a central concern in international education systems, there remains a need for context-sensitive, data-informed approaches that move beyond cross-sectional achievement measures. The study is guided by a value-added framework and contemporary models of school effectiveness, which emphasise growth-oriented assessment, instructional quality, and the role of home learning environments. Rather than relying solely on end-point achievement scores, the project examines literacy development longitudinally by administering equated initial and final literacy assessments. This approach allows for the estimation of student growth “beyond what might be expected,” accounting for prior achievement and contextual factors. The research addresses three core questions: The conceptual framework integrates multilevel models of learning with value-added analysis to capture the nested structure of educational data (students within classes, classes within schools). Student literacy growth is conceptualised as an outcome influenced by instructional practices, school organisation, and family-level support. The study also incorporates follow-up qualitative querying of selected classrooms to contextualise quantitative findings and explore pedagogical practices associated with higher-than-expected student growth. The international relevance of this research lies in its methodological contribution and its focus on literacy development within a rapidly reforming education system. By combining psychometrically robust literacy assessments, multilevel modelling, and value-added analysis, the study offers an evidence-informed framework applicable to other education systems seeking to strengthen literacy outcomes through growth-oriented evaluation rather than high-stakes accountability alone. Findings are expected to inform both policy and practice by highlighting effective pedagogical and organisational conditions for literacy development. Methodology, Methods, Research Instruments or Sources Used The study employs a longitudinal, quantitative research design complemented by targeted qualitative follow-up. Participants include students in Grades 3, 4, and 7 from 13 BI Education schools in Kazakhstan, with students nested within classes and schools. Literacy assessments are administered four times during the academic year, with equated initial and final tests used to estimate student growth in reading literacy. Literacy assessments are developed and validated using both Classical Test Theory and Modern Test Theory, including Rasch modelling. Item response data are analysed using the R programming language, with established psychometric packages used to estimate reliability, item functioning, and student ability. Common-item equating is applied to ensure comparability between initial and final assessments. In addition to student assessments, structured questionnaires are administered to parents and teachers. Parent surveys collect information on educational background, home literacy resources, reading practices, and parenting behaviours. Teacher surveys capture qualifications, pedagogical practices, instructional strategies, and classroom organisation. These instruments allow for the examination of both individual- and class-level predictors of literacy growth. Multilevel regression modelling is employed to analyse predictors of literacy development, accounting for the hierarchical structure of the data (students nested in classes, classes nested in schools). Final literacy ability estimates serve as dependent variables, with student-, teacher-, and school-level characteristics included as predictors. Value-added analysis is subsequently conducted to identify classrooms demonstrating high, moderate, and low contributions to student literacy growth. To contextualise quantitative findings, follow-up qualitative querying is conducted in selected classrooms. This non-invasive approach focuses on classroom culture, instructional practices, and student–teacher interactions, providing interpretive depth to the statistical results. Together, these methods offer a comprehensive and ethically grounded approach to understanding literacy development in school settings. Conclusions, Expected Outcomes or Findings The study is expected to identify key student-, teacher-, and family-level factors associated with literacy growth across primary and lower secondary grades. Anticipated findings include differential effects of pedagogical practices, classroom organisation, and home literacy environments on student growth, even after controlling for prior achievement. Value-added analysis is expected to reveal meaningful variation between classrooms, highlighting instructional contexts in which students demonstrate higher-than-expected literacy gains. These findings will support a shift from deficit-oriented interpretations of achievement toward growth-focused evaluation of teaching and learning. The integration of quantitative and qualitative evidence is expected to provide actionable insights into effective instructional practices and classroom conditions that support literacy development. The findings will be relevant for school leaders, policymakers, and researchers seeking to design professional development, curriculum support, and assessment systems aligned with evidence-informed improvement. More broadly, the study is expected to contribute to international discussions on literacy assessment, value-added modelling, and the ethical use of educational data to support student learning rather than punitive accountability. References OECD. (2013). Synergies for Better Learning: An International Perspective on Evaluation and Assessment. Paris: OECD Publishing. Ainscow, M., Chapman, C., & Hadfield, M. (2020). Changing education systems for equity and excellence. Educational Research, 62(3), 1–17. Costello, R., Elson, P., & Schacter, J. (2008). An introduction to value-added analysis. Journal of Catholic Education, 12(2), 195–212. Rasch, G. (1960). Probabilistic Models for Some Intelligence and Attainment Tests. Copenhagen: Danish Institute for Educational Research. Bates, D., Maechler, M., Bolker, B., & Walker, S. (2015). Fitting linear mixed-effects models using lme4. Journal of Statistical Software, 67(1), 1–48. | ||