Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
Please note that all times are shown in the time zone of the conference. The current conference time is: 19th Aug 2026, 21:29:09 EET
|
Daily Overview |
| Session | ||
99 ERC SES 04 I: Assessment and evaluation design and methodology
Paper Session
| ||
| Presentations | ||
99. Emerging Researchers' Group (for presentation at Emerging Researchers' Conference)
Paper Developing a Standards-Referenced Framework for Social Media Literacy: From Construct Definition to Rasch Validation Design The University of Sydney, Australia Presenting Author:Social media has become core infrastructure for accessing information, communicating, and participating in civic life. European policymakers have called for stronger media literacy to counter disinformation and support informed decision-making. Frameworks like DigComp 2.2 describe what competence looks like. But these frameworks cannot function as psychometrically validated measurement instruments—they describe competence but cannot locate individuals on a calibrated scale (Vuorikari et al., 2022). The lack of validated measurement creates difficulties. Evidence-informed curriculum development, assessment design, and cross-context proficiency comparison all require instruments that can establish where individuals stand. Substantial work on social media literacy (SML) has accumulated, yet the field remains conceptually fragmented (Cho et al., 2024; Valle et al., 2025). Most instruments rely on self-report measures. These capture what users perceive they can do, rather than what they actually demonstrate (Tandoc et al., 2021). Some newer performance measures move beyond self-report using IRT-based approaches. This is progress. But these tools typically target specific populations and do not address standards-referenced interpretation across age groups (Drake et al., 2023). Making defensible proficiency judgements across educational systems is difficult. We conducted a systematic scoping review and thematic synthesis of SML literature from 2015 to 2025. From this work, we propose a measurement framework applicable across ages and contexts. The framework operationalises SML through three components: Platform Awareness, Critical Evaluation, and Responsible Participation. These consolidate recurring domains in the literature. We describe proficiency using four levels. The four-level structure reflects theoretically distinct transitions in metacognitive regulation. L1 (Unaware) represents minimal recognition. L2 (Retrospective) involves post-hoc awareness—individuals identify relevant influences after engagement (Flavell, 1979). L3 (Real-time) marks a shift to in-process monitoring and deliberate adjustment during engagement (Schraw & Dennison, 1994). L4 (Habitual–Reflexive) describes internalised monitoring across contexts, including low-stakes use, where strategies transfer beyond single situations (Mezirow, 1997). Self-awareness runs through all levels as a difficulty factor rather than a separate component. Three questions guide the study: What are the core components and indicators of SML for standards-referenced measurement? How do we operationalise and design a rubric-based framework for Rasch calibration? What validity evidence is needed to support the intended uses of the measures? Our goal is translating competence language into a calibratable rubric that supports defensible proficiency judgements. The conference paper presents the three-component construct map, candidate indicators, draft performance level descriptors, and the validation workflow. We plan to validate the framework across cultural and educational contexts. Cross-cultural comparability will be addressed in later stages. The study contributes a validation logic for standards-referenced interpretation, supporting European efforts to assess SML comparably across populations and systems. Methodology, Methods, Research Instruments or Sources Used Psychometrically validated instruments for standards-referenced SML measurement are absent (Drake et al., 2023; Valle et al., 2025). We propose a Rasch-based workflow to address this gap, translating competence descriptions into calibrated proficiency scales for defensible cross-context interpretation. 1. Construct Definition through Systematic Synthesis (Completed) We conducted a systematic scoping review following PRISMA-ScR conventions (Tricco et al., 2018) of SML literature from 2015 to 2025. Twenty-five studies proposed competency frameworks or measurement dimensions. Thematic synthesis identified three recurring components: Platform Awareness (algorithmic curation, business logic, platform architecture), Critical Evaluation (source credibility, evidence-based claims, bias/framing analysis), and Responsible Participation (pre-posting and in-interaction responsibility). Construct mapping (Wilson, 2023) operationalised these into eight indicators with observable behaviours across four proficiency levels: L1 Unaware, L2 Retrospective, L3 Real-time, L4 Habitual–Reflexive. Self-awareness is embedded as a cross-cutting difficulty factor rather than a separate dimension. This is reflected in Performance Level Descriptors (PLDs) that discriminate adjacent levels, grounded in Flavell's (1979) awareness emergence and Schraw and Dennison's (1994) regulatory monitoring. 2. Expert Validation of Rubric Structure (Designed; Initial Implementation by Conference) Measurement specialists and SML domain experts (n=8–12) will evaluate construct coverage, indicator clarity, and PLD discriminability through structured focus groups. Qualitative feedback will inform descriptor refinement. Content Validity Index (CVI) ratings will quantify expert agreement on indicator relevance and level appropriateness, providing evidence of interpretive coherence prior to empirical calibration. 3. Rasch Calibration Workflow (Designed; Implementation Post-Conference) Scenario-based assessment tasks anchored to specific PLDs will undergo polytomous Rasch calibration if model fit is adequate. A field trial with adequate sample size for stable parameter estimation is planned. The validation workflow addresses four psychometric requirements: First, model specification compares the Rating Scale Model (RSM) with the Partial Credit Model (PCM) to determine optimal parameterisation. Second, Wright maps verify the theoretically predicted difficulty ordering across proficiency levels. Third, fit diagnostics using Infit/Outfit mean square (MNSQ) statistics identify misfitting items requiring revision. Fourth, Differential Item Functioning (DIF) tests examine whether demographic subgroups respond differently to items, establishing a foundation for later cross-context validation (e.g., Australia, Taiwan). 4. Conference Contribution We will present the construct map (three components, eight indicators) and draft PLDs, alongside interim evidence from expert consultation if available. We will also outline the pre-specified Rasch validation workflow for subsequent calibration. Conclusions, Expected Outcomes or Findings Three related outcomes emerge from this study, addressing the research questions and supporting standards-referenced SML assessment. 1. Framework specification. The construct map defines SML through three components grounded in theory: Platform Awareness, Critical Evaluation, and Responsible Participation. Eight indicators span four proficiency levels. Performance Level Descriptors treat metacognitive self-awareness as a cross-cutting difficulty factor. Post-hoc recognition appears at L2. L3 requires real-time monitoring. L4 captures habitual reflexivity. Recurring competency domains from across the literature are consolidated into a coherent measurement structure. This addresses conceptual fragmentation identified in recent reviews (Cho et al., 2024; Valle et al., 2025). 2. Validation evidence. Expert review provides content validity evidence. Measurement specialists and SML domain experts evaluate construct coverage, indicator clarity, and level discriminability through structured feedback. CVI ratings quantify how much experts agree on relevance and appropriateness. This strengthens interpretive coherence and helps determine readiness for field testing. 3. Calibration workflow specification. We present a fully specified Rasch validation workflow. Rubric-based competence descriptions are translated into calibrated measures. The workflow includes model comparison (RSM vs PCM), hierarchy evaluation using Wright maps to check alignment with expected difficulty ordering, fit diagnostics using Infit/Outfit MNSQ, and DIF checks across demographic groups for measurement invariance. This provides a replicable pathway. Descriptive frameworks can be converted into psychometrically sound instruments. Standards-referenced calibration has lacked these operational requirements. Beyond these methodological contributions, the study addresses a critical gap in EU digital literacy policy. While frameworks like DigComp 2.2 (Vuorikari et al., 2022) describe competence, they cannot support defensible proficiency judgements or cross-national comparison. The validated SML scale enables evidence-based curriculum placement and intervention targeting across European educational systems. It offers policymakers a measurement model that moves from descriptive taxonomies to calibrated assessment. References Andrich, D. (2004). Controversy and the Rasch model: a characteristic of incompatible paradigms? Med Care, 42(1 Suppl), I7-16. https://doi.org/10.1097/01.mlr.0000103528.48582.7c Cho, H., Cannon, J., Lopez, R., & Li, W. B. (2024). Social media literacy: A conceptual framework. New Media & Society, 26(2), 941-960, Article 14614448211068530. https://doi.org/10.1177/14614448211068530 Drake, A. P., Masur, P. K., Bazarova, N. N., Zou, W. T., & Whitlock, J. (2023). The youth social media literacy inventory: development and validation using item response theory in the US. Journal of Children and Media, 17(4), 467-487. https://doi.org/10.1080/17482798.2023.2230493 Flavell, J. H. (1979). Metacognition and Cognitive Monitoring: A New Area of Cognitive-Developmental Inquiry. American Psychologist, 34, 906-911. https://doi.org/https://doi.org/10.1037/0003-066X.34.10.906 Mezirow, J. (1997). Transformative Learning: Theory to Practice. New Directions for Adult and Continuing Education, 1997(74), 5-12. https://doi.org/https://doi.org/10.1002/ace.7401 Schraw, G., & Dennison, R. S. (1994). Assessing Metacognitive Awareness. Contemporary Educational Psychology, 19(4), 460-475. https://doi.org/https://doi.org/10.1006/ceps.1994.1033 Tandoc, E. C., Yee, A. Z. H., Ong, J., Lee, J. C. B., Xu, D., Han, Z., Matthew, C. C. H., Ng, J., Lim, C. M., Cheng, L. R. J., & Cayabyab, M. Y. (2021). Developing a Perceived Social Media Literacy Scale: Evidence from Singapore. International Journal of Communication, 15, 2484-2505. <Go to ISI>://WOS:000729944300009 Tricco, A. C., Lillie, E., Zarin, W., O'Brien, K. K., Colquhoun, H., Levac, D., Moher, D., Peters, M. D. J., Horsley, T., Weeks, L., Hempel, S., Akl, E. A., Chang, C., McGowan, J., Stewart, L., Hartling, L., Aldcroft, A., Wilson, M. G., Garritty, C.,…Straus, S. E. (2018). PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and Explanation. Ann Intern Med, 169(7), 467-473. https://doi.org/10.7326/m18-0850 Valle, N., Zhao, P. F., Freed, D., Gorton, K., Chapman, A. B., Shea, A. L., & Bazarova, N. N. (2025). Towards a Critical Framework of Social Media Literacy: A Systematic Literature Review. Review of Educational Research, 95(4), 701-746. https://doi.org/10.3102/00346543241247224 Vuorikari, R., Kluzer, S., & Punie, Y. (2022). DigComp 2.2: The Digital Competence Framework for Citizens – With new examples of knowledge, skills and attitudes (EUR 31006 EN; JRC128415). https://publications.jrc.ec.europa.eu/repository/handle/JRC128415 White, R., & Simpson, A. (Eds.). (2025). What’s the Evidence?: An Investigation into Teacher Quality. Routledge. https://doi.org/10.4324/9781003542575. Wilson, M. (2023). Constructing Measures: An Item Response Modeling Approach (2nd Edition ed.). Routledge. https://doi.org/https://doi.org/10.4324/9781003286929 99. Emerging Researchers' Group (for presentation at Emerging Researchers' Conference)
Paper Evaluating DIF in the Successful School Leadership Survey The University of Alabama, United States of America Presenting Author:The Successful School Leadership framework (SSL) is a model of best practices for school leadership, as validated through multiple studies over a twenty-year period (Leithwood et al., 2010; Leithwood & Jantzi, 2005; Sun & Leithwood, 2015). The Successful School Leadership survey was created as a measure of perceived school leadership quality under this framework. This survey is a 22-item Likert scale with 5-point response options ranging from “not at all confident” to “very confident”. This scale intends to measure four domains of leadership: (1) setting directions, (2) building relationships and developing people, (3) (re)designing the organization to support desired practices, and (4) improving the instructional program. Leithwood et al. (2023) provided strong validity evidence for this scale, reporting high reliability, measurement invariance for various person characteristics, and good person and item fit. This study also confirmed the four-factor structure and showed predictive validity evidence of high leadership scores for student achievement, even after controlling for relevant confounding variables. Leithwood et al. (2023) showed evidence of measurement invariance for gender and education level by using the Many Facet Rasch Model (MFRM) to measure the effect of these covariates at the group level, confirming that gender and education were not predictive of scale scores. However, this method of investigating measurement invariance does not consider differences in subgroup effects at the item level. This difference in item parameters across different groups of people when holding person abilities constant is known as differential item functioning (DIF) and is a common threat to measurement invariance and fairness in scale validation (Ackerman & Evans, 1994; Roussos & Stout, 1996). This study investigates the presence of DIF in the SSLS items for three covariates: gender, experience, and education level. We use traditional DIF detection methods evaluate these three covariates separately. However, recent interest in intersectional approaches to quantitative research (e.g., Albano et al., 2024; Bauer et al., 2021) underscores the importance of not just treating covariates as independent factors, but as more complex interdependent person characteristics. In response to this recent rise in interest surrounding intersectional approaches to quantitative research, we also use some newer, more exploratory DIF detection methods to investigate intersectional DIF in the SSLS items. The data used in this study were collected as part of the research conducted in Leithwood et al. (2023). All participants in this survey were teachers in two different neighboring school districts in the southeastern United States. After filtering participants based on missing data and willingness to participate, this dataset had a total of 136 respondents. Most of the sample population (83%) identified as women, and the rest of the population identified as men. 58% of the sample population held graduate degrees, and the rest held bachelor’s degrees. 38 participants had 1-3 years of experience, 29 participants had 4-10 years of experience, 20 had 11-15 years of experience, and 48 had 15 or more years of experience. Methodology, Methods, Research Instruments or Sources Used Before evaluating SSLS items for DIF, the Partial Credit Model (PCM) was fit to the data to assess basic properties of model-data fit (Masters, 1982). We used the TAM package in R (Robitzsch et al., 2023) to fit the PCM, which is specified as ln[(P_ni (x_i=k))/(P_ni (x_i=k-1) )]=θ_n-δ_i-τ_ik Where θ_n is person n’s ability level, δ_i is the difficulty of item i, and τ_ik is the rating scale category threshold for item i. Thresholds are located on the logit scale where the probability of response k and k-1 are equal. To assess model-data fit, we calculated infit and outfit for items. Many studies use widely accepted cut points, such as 0.8 to 1.2, as “critical ranges” for infit and outfit to determine if an item is misfitting. However, these general cut points range in value across different studies and are prone to high classification error (Wolfe, 2013). In this study, we use more robust empirically derived critical values for fit statistics via parametric bootstrapping, where bootstrapped 95% confidence intervals are used to determine critical ranges for infit and outfit (see Seol (2016) and Silvia Diaz et al. (2022) for a more in-depth discussion of this method). After checking model-data fit, we began assessing DIF using Lord’s chi-square method for DIF detection in the R package lordif (Choi et al., 2011; Lord, 1980). Lord’s Chi Square tests can check for both uniform and nonuniform DIF by testing for significant differences in models that do and do not contain DIF. We assessed DIF separately for gender, experience, and education level. We then used the Partial Credit Tree model in the R package psychotree (Komboz et al., 2018; Strobl et al., 2015) to assess intersectional DIF. The Partial Credit Tree model is a polytomous Rasch Tree model that uses recursive partitioning to iteratively search for DIF in combinations of covariates. This model allows us to look for intersectional DIF within each possible covariate combination, forgoing the need to predefine intersectional focal groups, making this method more exploratory than other similar intersectional approaches to DIF detection. Conclusions, Expected Outcomes or Findings The Partial Credit Model showed one item with disordered category thresholds (item 18), indicating that response categories for this item were inconsistent with the observed responses in this round of data collection. The fit analysis showed one item with substantial misfit for both infit and outfit based on the bootstrapped confidence intervals (item 8). The full paper will investigate the content of these items further to locate any potential sources of misfit. The chi-square tests for individual covariate DIF detection revealed DIF for gender and education level, but no DIF for experience. For gender, item 11 showed uniform DIF, indicating consistent differences in item parameters across the latent scale, and item 14 showed nonuniform DIF, revealing DIF that affected people differently at different locations on the latent scale. For Education, item 20 showed uniform DIF, and item 12 showed nonuniform DIF. The full paper will review these items qualitatively to locate any potential sources of item-level bias. However, the Partial Credit Tree showed no evidence of DIF for any covariate, returning a tree model with no node splits. This may indicate that the chi-square model had a level of Type I error in detecting DIF for four items across two different covariates. However, as discussed in Krist et al. (2025), imbalanced samples, like the one we have in this study, tend to lead to an underestimation of DIF and an increase in Type II error. This imbalanced sample size may also be the reason for the discrepancy in DIF detection between the two methods. This problem of imbalanced samples is somewhat unavoidable in the current context, given that most teachers in the US are women. Still, future research should explore this discrepancy further, possibly with more participants, more balanced samples, and different DIF detection models to verify DIF results. References Ackerman, T. A., & Evans, J. A. (1994). The influence of conditioning scores in performing DIF analyses. Applied Psychological Measurement, 18(4), 329–342. https://doi.org/10.1177/014662169401800404 Albano, T., French, B. F., & Vo, T. T. (2024). Traditional vs Intersectional DIF Analysis: Considerations and a Comparison Using State Testing Data. Applied Measurement in Education, 37(1), 57–70. https://doi.org/10.1080/08957347.2024.2311935 Bauer, G. R., Churchill, S. M., Mahendran, M., Walwyn, C., Lizotte, D., & Villa-Rueda, A. A. (2021). Intersectionality in quantitative research: A systematic review of its emergence and applications of theory and methods. SSM - Population Health, 14, 1–11. https://doi.org/10.1016/j.ssmph.2021.100798 Choi, S. W., Gibbons, L. E., & Crane, P. K. (2011). lordif: An R package for detecting Differential Item Functioning using iterative hybrid ordinal logistic regression/Item Response Theory and Monte Carlo simulations. Journal of Statistical Software, 39(8), 1-30. https://doi.org/10.18637/jss.v039.i08 Komboz, B., Strobl, C., & Zeileis, A. (2018). Tree-based global model tests for polytomous Rasch models. Educational and Psychological Measurement, 78(1), 128–166. https://doi.org/10.1177/0013164416664394 Krist, A. T., Wind, S. A., & Lakin, J. M. (2025). Rasch tree modeling for DIF analysis: How sample and item characteristics influence DIF detection. Applied Measurement in Education, 1–16. https://doi.org/10.1080/08957347.2025.2559811 Leithwood, K., Harris, A., & Strauss, T. (2010). Leading school turnaround: How successful leaders transform low-performing schools. San Francisco: Jossey-Bass. Leithwood, K., & Jantzi, D. (2005). A Review of Transformational School Leadership Research 1996-2005. Leadership and Policy in Schools, 4, 177-199. https://doi.org/10.1080/15700760500244769 Leithwood, K., Sun, J., Schumacker, R., & Hua, C. (2023). Psychometric properties of the successful school leadership survey. Journal of Educational Administration. https://doi.org/10.1108/JEA-08-2022-0115 Masters, G. N. (1982). A Rasch model for partial credit scoring. Psychometrika, 47(2), 149–174. https://doi.org/10.1007/BF02296272 Robitzsch, A., Kiefer, T., & Wu, M. (2023). TAM: Test analysis modules (R package version 4.0-2). https://CRAN.R-project.org/package=TAM Seol, H. (2016). Using the Bootstrap Method to Evaluate the Critical Range of Misfit for Polytomous Rasch Fit Statistics. Psychological Reports, 118(3), 937–956. https://doi.org/10.1177/0033294116649434 Silva Diaz, J. A., Köhler, C., & Hartig, J. (2022). Performance of Infit and Outfit Confidence Intervals Calculated via Parametric Bootstrapping. Applied Measurement in Education, 35(2), 116–132. Academic Search Premier. Strobl, C., Kopf, J., & Zeileis, A. (2015). Rasch trees: A new method for detecting differential item functioning in the Rasch model. Psychometrika, 80(2), 289-316. Sun, J., & Leithwood, K. (2015). Direction-setting school leadership practices: a meta-analytical review of evidence about their influence. School Effectiveness and School Improvement, 26(4), 499–523. https://doi.org/10.1080/09243453.2015.1005106 | ||
