Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
Please note that all times are shown in the time zone of the conference. The current conference time is: 19th Aug 2026, 21:29:42 EET
|
Daily Overview |
| Session | ||
09 SES 01 C: Meta-Analysis, Research Synthesis and Evidence Quality
Paper Session | ||
| Presentations | ||
09. Assessment, Evaluation, Testing and Measurement
Paper Working for Pay and Academic Achievement: A Cross-National Meta-Analysis of Gender-Moderated Effects Using PISA 2022 Data 1: Hacettepe University, Turkey (Türkiye); 2: Trakya University, Turkey (Türkiye) Presenting Author:Adolescents working in paid jobs may struggle with time management due to the very different expectations of work and school. When they add work obligations to their educational responsibilities, they face potential dangers that could negatively affect their academic performance (Greenhaus & Beutell, 1985). Recently, with economic instability and increasing social inequality becoming more apparent, a growing number of adolescents have been forced to combine their school education with paid work out of necessity rather than choice (Staff et al., 2010). Indeed, we cannot assert that a negative relationship always exists between student employment and academic achievement. Holding multiple social positions can indirectly positively influence success by contributing to the development of various social skills and self-efficacy (Sieber, 1974). The direction and magnitude of these effects depend on contextual factors such as work intensity, the nature of employment, and socioeconomic conditions (Marsh & Kleitman, 2005). Therefore, local structural conditions deeply influence the outcomes of student employment. This study examines how working in a paid job during adolescence shapes academic success in different national contexts, focusing on gender dimensions. Since cultural norms regarding adolescent employment and gender expectations vary significantly across societies, understanding gender-related patterns necessitates a cross-national comparative approach. The study addresses three interconnected research questions:
This study critically examines how large-scale international assessments that shape education policies can shed light on the complex interaction between the economy and education. By disaggregating the effects across 79 countries, the heterogeneity of the findings was examined, yielding results resistant to the homogenizing tendencies inherent in global assessment frameworks (Rutkowski & Rutkowski, 2016). The marked cross-country differences observed in the results indicate the importance of context-sensitive approaches rather than uniform policy recommendations. Furthermore, this research has addressed some of the crises facing contemporary education systems. Recent global crises have increased financial pressure on families, leading more adolescents to enter paid employment out of necessity rather than choice (Lee & Staff, 2007). If the outcomes of combining work and school are truly context-dependent, understanding which structural conditions exacerbate or mitigate negative effects becomes an urgent policy issue. The findings have important implications for labor regulations designed to protect adolescents' educational opportunities in various national contexts, as well as for countries' education policies and social support systems (Post & Pong, 2000). Methodology, Methods, Research Instruments or Sources Used This study employed a two-stage analytical design that combined multilevel modeling with meta-analysis, using data from the 2022 cycle of the Programme for International Student Assessment (PISA) (OECD, 2024). The sample comprised 79 countries and economies participating in PISA 2022, representing diverse geographic, economic, and cultural contexts. Student achievement in mathematics, reading, and science was measured using five plausible values per domain, as provided by PISA. Students’ engagement in paid work was assessed using the WORKPAY index, which captures how frequently students work for pay before or after school during a typical school week. Gender was included both as a covariate and as a moderator. All continuous predictors were standardized to facilitate cross-country comparability. The analytical procedure consisted of two stages. In the first stage, two-level hierarchical linear models were estimated separately for each country, with students nested within schools (Raudenbush & Bryk, 2002). These models included main effects of paid work and gender, as well as interaction terms to test gender moderation. Random intercepts at the school level were specified to account for clustering. Each model was estimated separately for ten plausible values, and results were combined using Rubin’s (1987) multiple imputation rules. Student- and school-level sampling weights were applied to account for PISA’s complex sampling design (Rutkowski et al., 2010). In the second stage, country-specific estimates obtained from the multilevel models were synthesized using random-effects meta-analysis (Borenstein et al., 2009). Separate meta-analyses were conducted for each academic domain and predictor, yielding a total of nine meta-analytic models. Cross-national heterogeneity was evaluated using Cochran’s Q statistic, the I² index (Higgins & Thompson, 2002), and the between-study variance component τ². These indicators allowed for an assessment of whether observed variation in effect sizes reflected substantive differences between national contexts rather than sampling variability. This two-stage approach preserves the multilevel structure of the data while enabling the examination of cross-national variability in effects (Snijders & Bosker, 2012). All analyses were conducted in R (R Core Team, 2024). Multilevel models were estimated using the mixPV function (Huang, 2024), which builds on the WeMix package (Bailey et al., 2023) and supports weighted estimation with plausible values. Meta-analyses were performed using the metafor package (Viechtbauer, 2010). Conclusions, Expected Outcomes or Findings Meta-analytic findings show a consistent negative relationship between paid work and academic achievement across reading, science, and mathematics. Students who work more frequently perform worse in reading (β = −0.152, p < .001), science (β = −0.143, p < .001), and mathematics (β = −0.121, p < .001). These findings suggest that paid work interferes with students’ academic responsibilities, supporting role conflict theory (Greenhaus & Beutell, 1985). Despite this overall pattern, substantial heterogeneity is observed across all domains, with I² values exceeding 89%, indicating considerable cross-national variation in the effects of paid work on achievement (Higgins & Thompson, 2002). Geographic patterns reveal that economically developed countries—such as Scandinavian nations, Canada, the United States, and Australia—exhibit stronger negative associations. Although adolescent employment in these contexts is often not economically necessary, education systems typically assume that students devote most of their time and energy to schooling. Consequently, paid work may conflict more strongly with academic expectations, leading to greater performance declines (Post & Pong, 2000). Gender differences follow established patterns: male students perform better in mathematics and science, whereas female students outperform males in reading (OECD, 2023). Interaction effects examining whether the impact of paid work varies by gender are statistically significant but small (β = −0.007 to −0.018), suggesting that the challenge of balancing work and school affects male and female students in a largely similar manner (Voydanoff, 2005). Overall, these findings indicate that policies targeting working students should be context-sensitive rather than uniform, as the academic risks associated with paid work differ markedly across countries, particularly in developed economies. References Bailey, P., Kelley, C., Nguyen, T., & Huo, H. (2023). WeMix: Weighted mixed-effects models using multilevel pseudo maximum likelihood estimation (R package version 4.0.1). https://CRAN.R-project.org/package=WeMix Borenstein, M., Hedges, L. V., Higgins, J. P. T., & Rothstein, H. R. (2009). Introduction to meta-analysis. Wiley. Greenhaus, J. H., & Beutell, N. J. (1985). Sources of conflict between work and family roles. The Academy of Management Review, 10(1), 76–88. https://doi.org/10.2307/258214 Higgins, J. P. T., & Thompson, S. G. (2002). Quantifying heterogeneity in a meta-analysis. Statistics in Medicine, 21(11), 1539–1558. https://doi.org/10.1002/sim.1186 Huang, F. L. (2024). Using plausible values when fitting multilevel models with large-scale assessment data using R. Large-Scale Assessments in Education, 12(1). https://doi.org/10.1186/s40536-024-00192-0 Lee, J. C., & Staff, J. (2007). When work matters: The varying impact of work intensity on high school dropout. Sociology of Education, 80(2), 158-178. https://doi.org/10.1177/003804070708000204 Marsh, H. W., & Kleitman, S. (2005). Consequences of employment during high school: Character building, subversion of academic coals, or a threshold? American Educational Research Journal, 42(2), 331-369. https://doi.org/10.3102/00028312042002331 OECD (2023), PISA 2022 Results (Volume I): The State of Learning and Equity in Education, PISA, OECD Publishing, Paris, https://doi.org/10.1787/53f23881-en. OECD (2024), PISA 2022 Technical Report, PISA, OECD Publishing, Paris, https://doi.org/10.1787/01820d6d-en. Post, D., & Pong, S. (2000). Employment during middle school: The effects on academic achievement in the U.S. and abroad. Educational Evaluation and Policy Analysis, 22(3), 273-298. https://doi.org/10.3102/01623737022003273 Raudenbush, S. W., & Bryk, A. S. (2002). Hierarchical linear models: Applications and data analysis methods (2nd ed.). Sage. Rubin, D. B. (1987). Multiple imputation for nonresponse in surveys. Wiley. https://doi.org/10.1002/9780470316696 Rutkowski, L., Gonzalez, E., Joncas, M., & von Davier, M. (2010). International large-scale assessment data: Issues in secondary analysis and reporting. Educational Researcher, 39(2), 142-151. https://doi.org/10.3102/0013189X10363170 Rutkowski, L., & Rutkowski, D. (2016). A call for a more measured approach to reporting and interpreting PISA results. Educational Researcher, 45(4), 252-257. https://doi.org/10.3102/0013189X16649961 Sieber, S. D. (1974). Toward a theory of role accumulation. American Sociological Review, 39(4), 567–578. https://doi.org/10.2307/2094422 Snijders, T. A. B., & Bosker, R. J. (2012). Multilevel analysis: An introduction to basic and advanced multilevel modeling (2nd ed.). Sage. Staff, J., Schulenberg, J. E., & Bachman, J. G. (2010). Adolescent work intensity, school performance, and Academic Engagement. Sociology of Education, 83(3), 183-200. https://doi.org/10.1177/0038040710374585 Viechtbauer, W. (2010). Conducting meta-analyses in R with the metafor package. Journal of Statistical Software, 36(3), 1–48. https://doi.org/10.18637/jss.v036.i03 09. Assessment, Evaluation, Testing and Measurement
Paper Seductive Details, Cognitive Load, and Academic Achievements: A Multi-level Meta-analysis and MASEM east china normal university Presenting Author:Introduction The inclusion of visually engaging but instructionally irrelevant elements—termed seductive details (Harp & Mayer, 1998)—has sparked ongoing debate regarding their pedagogical value. While such features may enhance engagement by capturing attention and stimulating situational interest, they risk diverting cognitive resources from core content, potentially impairing learning. This trade-off poses a significant challenge in designing digital instructional materials, where balancing engagement and cognitive load is essential. In response, the present meta-analysis synthesizes two decades of research on the seductive details effect, aiming to clarify under what conditions these elements hinder or support learning and to inform evidence-based design practices for effective and engaging instruction. Literature review Seductive details effect and mediating role of cognitive load A central question in seductive details research is whether such elements help or hinder learning. Early studies typically reported negative effects, particularly for learners with lower working memory who struggle to filter irrelevant content (Harp & Mayer, 1997). More recent findings, however, are mixed. Some suggest seductive details enhance situational interest without improving deep learning or metacognition (Magner et al., 2014; Jager & Wiley, 2014), while others propose they act as cognitive breaks that help sustain attention and reduce fatigue (Fries et al., 2019; Szpunar et al., 2013). These inconsistencies underscore the need to explore boundary conditions such as learner characteristics, task demands, and design features. The seductive details effect is grounded in Cognitive Load Theory (CLT; Sweller et al., 1998) and the Cognitive Theory of Multimedia Learning (CTML; Mayer, 2014), which differentiate intrinsic, extraneous, and germane load. Seductive details are thought to increase extraneous load and impaired recall performance (Wesenberg et al., 2025), though the mediating effect of extraneous cognitive load is not consistently observed (Colliot & Boucheix, 2024). Moderators variables Recent research has moved beyond confirming the seductive details effect to investigating its boundary conditions. This reflects a broader recognition that instructional outcomes depend on learner characteristics and environmental context. Moderating factors include age (Scherer et al., 2023), prior knowledge (Magner et al., 2014), working memory capacity (Sanchez & Wiley, 2006), and motivation (Wang & Adesope, 2016). Contextual elements like instructional techniques (Eitel et al., 2019; Wang et al., 2017) and learning environments (Bender et al., 2021; Fries et al., 2019) also influence the effect. Emerging evidence further suggests that interactions between learner traits and seductive detail types shape outcomes (Wesenberg et al., 2024). Given these inconsistencies, a meta-analytic approach is needed to clarify when and for whom seductive details help or hinder learning (Mayer, 2019). Current meta-analysis Two prior meta-analyses examined the effects of seductive details on academic achievement. Rey (2012) reported small-to-medium negative effects on recall (d = -0.30) and transfer (d = -0.48), while Sundararajan and Adesope (2020) analyzed 68 effect sizes from 42 studies and found a similar effect (g = -0.33). However, these studies did not address effect size non-independence or test cognitive load as a mediator. Since 2020, the number of relevant studies has nearly doubled, highlighting the need for an updated analysis. This meta-analysis includes studies published between 2019 and 2025 and adopts a multilevel approach to account for dependencies among effect sizes. The study is guided by four research questions: RQ1: What is the overall effect of seductive details on academic achievement? RQ2: How do they affect specific learning outcomes? RQ3: What factors moderate these effects? RQ4: To what extent does cognitive load mediate the relationship between seductive details and learning outcomes? Methodology, Methods, Research Instruments or Sources Used Selection criteria The selected literatures should met the inclusion criteria which is reported in Table 1. Literature search and selection. As shown in Figure 1, the literature search followed the PRISMA guidelines (Page et al., 2021; Moher et al., 2009). We conducted a literature search in the electronic databases including Web of Science, PsycINFO, Scopus, ProQuest, and ERIC (Education Resources Information Center) using the search syntax shown in Table 2. Consequently, a total of 51 publications were included in this review at the end. Literature coding Before coding, the authors held detailed discussions to finalize the coding criteria. A random sample of 10 references (20%) was independently coded to ensure consistency, with disagreements resolved through discussion. The remaining 41 references were coded using the agreed scheme. Inter-coder reliability was high (Cohen’s kappa = 0.91–0.96, p < 0.001). Following Wilson and Lipsey (2001), six coding schemes were applied (see Table 3), focusing on effect sizes and study-level variables. Effect sizes were categorized by learning outcome (recall, comprehension, transfer), and 14 study-level variables were coded to support moderation and mediation analyses. Data analysis All of the analyses, including overall effect size calculation, sensitivity analysis, moderation analysis, and mediation analysis, were conducted in R Studio (R Core Team, 2017) with the metafor package (Viechtbauer, 2010) and metaSEM package (Cheung, 2015; Cheung & Cheung, 2016). In this study, the effect size represents the effects of seductive details on the academic achievements. We used Hedges’ g (Hedges, 1982) as the standardized mean difference effect size metric. For the mediating analysis, we employed a two-stage meta-analytic structural equation modeling (MASEM) approach (Cheung & Chan, 2005), using 16 effect sizes from eight studies—including four derived from raw data obtained by contacting the authors of two studies (Wang et al., 2019; 2021). Conclusions, Expected Outcomes or Findings The seductive details effect (RQ1 & RQ2) The present meta-analysis confirms that seductive details have a small but significant negative effect on academic achievement (g = -0.16), aligning with earlier studies (Rey, 2012; Sundararajan & Adesope, 2020) but reporting a smaller effect. This difference likely stems from methodological improvements: our multi-level model accounted for data dependencies which were overlooked in prior pairwise meta-analyses. Across achievement types, seductive details negatively affected recall (g = -0.17), comprehension (g = -0.19), and transfer (g = -0.12). While these outcomes vary in cognitive complexity—from lower-order (recall) to higher-order (comprehension and transfer) cognitive abilities—the differences in effect size were minimal (maximum Δg = 0.07). This suggests that seductive details consistently impair learning across cognitive levels, with no strong evidence that higher-order tasks are more affected than lower-order ones. The moderating effects (RQ3) Three significant moderators emerged. First, sample size influenced effects: small (N ≤ 50, g = -0.49) and large samples (N ≥ 100, g = -0.16) showed significant negative effects, while medium-sized samples did not, possibly due to statistical power issues or bias. Second, text language mattered: English (g = -0.23) and German (g = -0.14) texts produced negative effects, while Chinese texts showed no significant impact—likely due to limited data and cultural familiarity with seductive content in English. Third, applied environment moderated outcomes: seductive details in technology-based contexts (g = -0.19) impaired learning significantly, whereas paper-and-pen settings had no effect, suggesting that digital environments may amplify distraction and cognitive overload, while traditional formats promote sustained focus. The mediating effects (RQ4) The mediation analysis supports CTML by showing that seductive details increase extraneous cognitive load, which in turn hinders academic achievement. Both direct and indirect effects were significant, suggesting other mediators like situational interest. Instructionally, designers should balance engagement with cognitive efficiency, minimizing seductive elements—especially for novices or complex topics—to reduce load and enhance deep learning. References Cheung, M. W. L. (2015). metaSEM: An R package for meta-analysis using structural equation modeling. Frontiers in Psychology, 5. https://doi.org/10.3389/fpsyg.2014.01521 Cheung, M. W. L., & Chan, W. (2005). Meta-analytic structural equation modeling: A two-stage approach. Psychological Methods, 10(1), 40–64. https://doi.org/10.1037/1082-989x.10.1.40 Cheung, M. W. L., & Cheung, S. F. (2016). Random‐effects models for Meta‐Analytic Structural Equation Modeling: Review, issues, and illustrations. Research Synthesis Methods, 7(2), 140–155. https://doi.org/10.1002/jrsm.1166 Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Lawrence Erlbaum Associates. Colliot, T., & Boucheix, J. (2024). Exploring the effects of seductive details and illustration dynamics on young children’s performance in an origami task. Journal of Computer Assisted Learning, 40, 437-451. https://doi.org/10.1111/jcal.12879 Fries, L., DeCaro, M. S., & Ramirez, G.(2019). The lure of seductive details during lecture learning. Journal of Educational Psychology, 111, 736-749. http://dx.doi.org/10.1037/edu0000301 Harp, S., & Mayer, R. E. (1997). The role of interest in learning from scientific text and illustrations: On the distinction between emotional interest and cognitive interest. Journal of Educational Psychology, 89, 92–102. https://doi.org/10.1037/0022-0663.89.1.92 Harp, S. F., & Mayer, R. E. (1998). How seductive details do their damage: A theory of cognitive interest in science learning. Journal of Educational Psychology, 90, 434–441. https://doi.org/10.1037/0022-0663.90.3.414 Hedges, L. V. (1982). Estimation of effect size from a series of independent experiments. Psychological Bulletin, 92(2), 490–499. https://doi.org/10.1037/0033-2909.92.2.490 Mayer, R. E. (2014). Cognitive theory of multimedia learning. In R. E. Mayer (Ed.), The Cambridge handbook of multimedia learning (2nd ed., pp. 43-71). Cambridge University Press. https://doi.org/10.1017/CBO9781139547369.005 Mayer, R. E. (2019). Taking a new look at seductive details. Applied Cognitive Psychology, 33, 139-141. https://doi.org/10.1002/acp.3503 Mayer, R. E. (2021). Multimedia learning (3rd ed.). Cambridge University Press. Moher, D., Liberati, A., Tetzlaff, J., & Altman, D. G. (2009). Preferred reporting items for systematic reviews and meta-analyses: The Prisma statement. Journal of Clinical Epidemiology, 62(10), 1006–1012. https://doi.org/10.1016/j.jclinepi.2009.06.005 Page, M. J., McKenzie, J. E., Bossuyt, P. M., Boutron, I., Hoffmann, T. C., Mulrow, C. D., Shamseer, L., Tetzlaff, J. M., Akl, E. A., Brennan, S. E., Chou, R., Glanville, J., Grimshaw, J. M., Hróbjartsson, A., Lalu, M. M., Li, T., Loder, E. W., Mayo-Wilson, E., McDonald, S., … Moher, D. (2021). The Prisma 2020 statement: An updated guideline for reporting systematic reviews. International Journal of Surgery, 88, 105906. https://doi.org/10.1016/j.ijsu.2021.105906 09. Assessment, Evaluation, Testing and Measurement
Paper Impact of Portfolio Assessment on Student Achievements and Attitudes in Türkiye: A Meta-Analysis Study TED University, Turkey (Türkiye) Presenting Author:Portfolio assessment, which can be defined as a purposeful collection of a student’s work that tells a story of the student’s efforts, progress, or achievement in one or more areas (Arter & Spandel, 1992), has received scholarly attention. Depending on various purposes, Miller et al. (2009) summarize portfolio use in four dimensions as (a) instruction and assessment, (b) current accomplishments and progress, (c) showcase and documentation, and (d) finished and working (p. 293), which do not only show the instructional and evaluative value of portfolio use, but also highlight the formative contributions to self-reflection, higher-order thinking, and self-regulation. Either paper-based or electronic, the use of portfolios has been shown to improve several variables in related research. The redundancy of findings specifically related to the effects of portfolio assessment on student achievements and attitudes, as the most frequently examined dependent variables, suggests a systematic review not only to interpret the magnitude of influence in terms of effect size, but also to facilitate scholars in designing their future studies. The following research questions, with achievement and attitude as the dependent variables and portfolio, either paper-based or electronic, as the independent variable, were addressed in this study: (1) What is the influence of portfolio use on student achievement? (2) What is the influence of portfolio use on student attitude? (3) How do several moderator variables (year of publication, type of publication, education level, subject, and type of portfolio) influence the effects of portfolio use? One novel contribution of the present meta-analysis is that it provides a review of extant research examining the effect of portfolio use on achievement and attitude in the context of Türkiye by including more studies proposed by search of related database and it quantifies the findings related to several moderators. Methodology, Methods, Research Instruments or Sources Used This research used meta-analysis to address questions pertaining to effects of portfolio assessment on student achievements and attitudes in scholarly work in Türkiye. For research purposes, experimental or a quasi-experimental research examining the effects of portfolios assessment on student achievement and attitude were included in analyses. The rigorous search of the premium databases in Turkiye yielded a total of 62 studies with 62 effect sizes on achievement (N=3.664) and 30 effect sizes on attitudes (N=1.779). Hedges’ g, a less biased measure of effect size (Lambert & Alhassoon, 2015), was used to determine the extent of the difference between achievement and attitude scores of experimental and control groups. As suggested by Field and Gillett (2010) for studies of meta-analysis in social sciences, random-effects model was employed for the analyses. Computations were carried out using Comprehensive Meta-Analysis Version 4 (Borenstein et al., 2022). The robust review of related research, the use of inclusion and exclusion criteria, and the reporting of effect sizes ensured the validity of the study. The reliability was provided with the agreement between two coders with a doctoral degree in education, who carried out the selection of studies, analysis and interpretation of data. Several statistics were employed to test the data for heterogeneity and publication bias (Borenstein et al., 2021). I2 and Q statistics were performed in the measurement of heterogeneity. The significance level of Q value (e.g. p<.01) and I2 values of 25%, 50% and 75% that represent low, moderate and high heterogeneity, respectively (Higgins & Thompson, 2002) were used to detect heterogeneity. Publication bias was examined with Egger’s Regression Test, Rosenthal’s FSN, and Duval & Tweedie’s Trim and Fill (TFM). The symmetry observed in the funnel plots and the nonsignificant p values (>.05) in Egger’s Regression Test would show the nonexistence of publication bias. For the Rosenthal’s FSN, the 5k+10 value (Fragkos et al., 2014; Mullen et al., 2001) was used as a cut-off point to observe the nonexistence of publication bias. Analog to ANOVA was used to explore the moderating effect of year of publication, type of publication, education level, subject, and type of portfolio on mean effect sizes in studies examining the effect of portfolio assessment on both achievement and attitude. Conclusions, Expected Outcomes or Findings The random effects model reflected the mean effect size as 0,998 (SE = 0,11, 95% CI = [0,79, 1,21]), which implied a large positive effect of portfolio assessment on students’ achievements; whereas the mean effect size of 0,676 (SE = 0,10, 95% CI = [0,48, 0,88]) for studies examining the influence of portfolio assessment on student attitudes implied a medium positive effect (Cohen, 1988). Further analyses for both achievement and attitude showed that data were highly heterogeneous and displayed no clear evidence of publication bias. A cumulative analysis of all the Analog to ANOVA results regarding the moderators in both achievement and attitude studies revealed the significant effect of subject only. In other words, subject as a variable was found to be the source of variance moderating the magnitude of effect sizes significantly in both achievement and attitude studies. References Arter, J., & Spandel, V. “Using Portfolios of Student Work in Instruction and Assessment.” Educational Measurement: Issues and Practice, 1992, 11, 36–44. Borenstein, M., Hedges, L. V., Higgins, J. P. T., & Rothstein, H. R. (2021). Introduction to Meta-Analysis (Second ed.). Wiley. Borenstein, M., Hedges, L. E., Higgins, J. P. T., & Rothstein, H. R. (2022). Comprehensive Meta-Analysis Version 4. In Biostat, Inc. www.Meta-Analysis.com Field, A. P., & Gillett, R. (2010). How to do a meta-analysis. British Journal of Mathematical and Statistical Psychology, 63(3), 665-694. https://doi.org/10.1348/000711010X502733 Fragkos, K. C., Tsagris, M., & Frangos, C. C. (2014). Publication bias in meta-analysis: confidence intervals for Rosenthal’s fail-safe number. International scholarly research notices, 2014. https://doi.org/10.1155/2014/825383 Higgins, J., & Thompson, S. G. (2002). Quantifying heterogeneity in a meta-analysis. Statistics in Medicine, 21(11), 1539-1558. Lambert, J.E., & Alhassoon, O.M. (2015) Trauma-focused therapy for refugees: Meta-analytic findings. Journal of Counseling Psychology, 62(1), 28–37. https://doi.org/10.1037/cou0000048 Miller, M.D., Linn, R.L., & Gronlund, N.E. (2009). Measurement and assessment in teaching (10th ed). Pearson Education. Mullen, B., Muellerleile, P., & Bryant, B. (2001). Cumulative meta-analysis: a consideration of indicators of sufficiency and stability. Personality and Social Psychology Bulletin, 27(11), 1450. 09. Assessment, Evaluation, Testing and Measurement
Paper When Students Do Not Give Good Effort: A Research Synthesis with Implications for Validity 1: Umeå University, Sweden; 2: University of Cyprus, Cyprus Presenting Author:Across classroom, national, and international assessments, score interpretations typically assume that examinees invest sufficient effort. Yet a substantial body of evidence shows that test‑taking effort varies systematically with stakes, task characteristics, student beliefs, and context, with non‑trivial effects on scores and measurement properties. Based on a synthesis of our own research as well as the research of others, the present paper argues that assessment outcomes are a joint function of skill and will and that validity arguments and score uses should explicitly model, measure, and communicate examinee effort. Theoretically, expectancy–value theory provides an overarching account of why students choose to invest effort (expectancies and values, including attainment, intrinsic, utility, and cost), while item-level models such as the demands–capacity framework and the rapid-guessing account explain how engagement fluctuates across items within a session. Empirical research demonstrates that effort varies by stakes, item type and position, task demands, and student characteristics (e.g., beliefs, ability), and that both self-reported and behavioral indicators predict performance; meta-analytic evidence shows moderate to strong associations, while relationships between different measures of effort on the other hand tend to be low. Computer-based assessments enable high-resolution behavioral indicators (e.g., response-time effort and action counts) that identify disengagement (rapid guessing, perfunctory CR responses), though threshold selection and misclassification remain concerns. Insufficient effort poses construct-irrelevant variance that can inflate item difficulty and discrimination, attenuate reliability, and produce misleading correlations with external variables; effects are especially salient for low-stakes tests, longer sessions, and constructed-response items. International assessment evidence suggests that while aggregate rankings may be relatively robust to filtering disengaged responses, individual scores are affected and program decisions can be distorted when effort is ignored. We conclude that effort should be assessed whenever feasible (via brief self-report and process data), explicitly considered in validity arguments and score-use decisions, and, where warranted, incorporated into modeling (e.g., effort-moderated IRT) or reporting (effort annotations in formative contexts). Future research should strengthen theory linking motivation, emotions (e.g., boredom), and process data, refine probabilistic engagement indices that allow partial effort, and evaluate equity implications across groups and cultures. Methodology, Methods, Research Instruments or Sources Used This paper employs a theory-informed review methodology synthesizing conceptual frameworks, methodological approaches, and empirical findings on test-taking effort across classroom, national, and international assessments. First, we map the construct by positioning test-taking effort as the outcome of a motivational state tied to an assessment task and define it as the persistent investment of cognitive resources with the goal of accurately representing what one knows and can do. We adopt expectancy–value theory as a macro-level framework for antecedents of effort (expectancy for success; attainment, intrinsic, utility value; perceived costs) and integrate item-level process accounts—the demands–capacity model and rapid-guessing theory—to explain within-test dynamics of engagement. Second, we review measurement strategies: (a) self-report instruments (e.g., Student Opinion Scale; expectancy–value questionnaires; effort thermometers) used primarily in paper-based or hybrid settings, discussing their advantages (ease, global coverage) and limitations (bias, accuracy); and (b) behavioral indicators enabled by computer-based testing, notably item response times, counts of actions/keystrokes, and missingness, which support indices such as Response-Time Effort (RTE) and partial-engagement measures. We summarize thresholding approaches for distinguishing solution behavior from rapid guessing, highlight misclassification risks, and cover extensions to constructed-response items (e.g., short perfunctory responses, idle responding). Third, we examine contextual, item, and person correlates of effort (stakes, length, difficulty, item position; subject domain; gender, ability, socio-economic and cultural factors) and consequences for psychometrics (item parameter bias, DIF, reliability) and score interpretations (formative feedback, pilot studies, national evaluations, large-scale comparisons). Finally, we articulate implications for validity arguments and score use by proposing that evidence about engagement (distributions of RTE, item-position trends, CR/MC differentials) become part of routine program monitoring and reporting, with feasibility considerations noted for short tests, small samples, and low prevalence of disengagement. Conclusions, Expected Outcomes or Findings Our synthesis converges on a simple but consequential insight: assessment scores reflect both what students know and what they are willing to demonstrate under specific conditions. When the assumption of sufficient effort is violated—particularly in low-stakes contexts, longer sessions, and demanding or late-position items—construct-irrelevant variance is introduced that can compromise item calibration, distort reliability and validity evidence, and mislead classroom and system-level decisions. At the same time, maximal effort may not always be necessary; the sufficiency of engagement depends on the intended interpretation and use. Accordingly, programs and researchers should: (1) monitor effort routinely using brief self-report and process data; (2) report effort-aware indicators alongside scores for formative uses; (3) consider analytic adjustments such as motivation filtering or effort-moderated models where decisions warrant it; and (4) design assessments with engagement in mind (balanced item types, strategic breaks, clear relevance cues). Advancing theory that connects expectancy–value appraisals with item-level behaviors, emotions like boredom, and performance will sharpen predictions about when and for whom disengagement arises. With continued refinement of probabilistic engagement indices and practical operational guidelines, integrating effort into validity arguments can render interpretations fairer, more accurate, and better aligned with the realities of modern assessment. References Eklöf, H. (2010). Skill and will: Test-taking motivation and assessment quality. Assessment in Education, 17(4), 345–356. Eklöf, H., & Knekta, E. (2017). Using large-scale educational data to test motivation theories. International Journal of Quantitative Research in Education, 4(1–2), 52–71. Goldhammer, F., Martens, T., & Lüdtke, O. (2017). Conditioning factors of test-taking engagement in PIAAC. Large-scale Assessments in Education, 5(18), 1–25. Lindner, M. A., Lüdtke, O., & Nagy, G. (2019). The onset of rapid-guessing behavior over the course of testing time. Frontiers in Psychology, 10, 1533. Michaelides, M. P., & Ivanova, M. (2022). Response time as an indicator of test-taking effort in PISA. Psychological Test and Assessment Modeling, 64(3), 304–338. Michaelides, M. P., Ivanova, M., & Avraam, D. (2024). The impact of filtering out rapid-guessing examinees on PISA 2015 country rankings. Psychological Test and Assessment Modeling, 66(1), 50–62. Rios, J. (2021). Improving test-taking effort in low-stakes group-based educational testing: A meta-analysis. Applied Measurement in Education, 34(2), 85–106. Silm, G., Pedaste, M., & Täht, K. (2020). The relationship between performance and test-taking effort: A meta-analytic review. Educational Research Review, 31, 100335. Wise, S. L. (2017). Rapid-guessing behavior: Its identification, interpretation, and implications. Educational Measurement: Issues and Practice, 36(4), 52–61. Wise, S. L., & DeMars, C. E. (2005). Low examinee effort in low-stakes assessment: Problems and potential solutions. Educational Assessment, 10(1), 1–17. Wise, S. L., & DeMars, C. E. (2006). An application of item response time: The effort-moderated IRT model. Journal of Educational Measurement, 43(1), 19–38. Wise, S. L., & Smith, L. F. (2016). The validity of assessment when students don’t give good effort. In G. T. L. Brown & L. R. Harris (Eds.), Handbook of human and social conditions in assessment (pp. 204–220). Routledge. | ||
