Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
Please note that all times are shown in the time zone of the conference. The current conference time is: 19th Aug 2026, 21:30:06 EET
|
Daily Overview |
| Session | ||
09 SES 16 C: Higher Education and Subject-Specific Exam Assessment
Paper Session | ||
| Presentations | ||
09. Assessment, Evaluation, Testing and Measurement
Paper Learning-Effective Design of Formative Peer and Self-Assessment and Profession-Related Conditions for Success: Findings from an Intervention Study 1: University of Bern; 2: University of Teacher Education Lucerne, Switzerland Presenting Author:Formative assessment (FA), also known as assessment for learning, is understood as an instructional practice in which evidence of student learning is continuously elicited, interpreted, and used to adapt teaching and learning to students’ needs (Cauley & McMillan, 2010). By supporting ongoing instructional adjustment and addressing diverse learning needs, FA has become a central pedagogical approach for promoting inclusive, learning-oriented classroom practices. Within this framework, formative peer and self-assessment (PASA) are regarded as particularly powerful practices, as they actively involve students in assessment processes and extend opportunities for feedback and reflection beyond teacher-led assessment (Black & Wiliam, 2018). Numerous studies demonstrate that FA enhances both cognitive and noncognitive outcomes, including academic achievement, motivation, self-efficacy, and self-regulation (e.g. Granberg et al., 2021; Rakoczy et al., 2019; Yan et al., 2022). However, evidence from mathematics education also points to mixed effects of FA and PASA on student learning outcomes, which are commonly attributed to variation in implementation quality (Maskos et al., 2025). These findings highlight the pivotal role of teachers, as the learning effectiveness of FA depends largely on how assessment practices are planned, enacted, and used within classroom instruction. This challenge is further reflected in observational research. FA is often only partially and inconsistently implemented in classroom practice and frequently with limited instructional quality (Buholzer et al., 2020; Gotwals et al., 2015), particularly with regard to PASA. Accordingly, teachers play a central role in FA by interpreting evidence of student learning and using this information to adapt instruction and provide feedback that supports further learning (Narciss et al., 2020). In the context of PASA, professional demands increase further, as teachers are required to design suitable tasks, establish transparent criteria, scaffold assessment processes, and orchestrate classroom interaction (Wanner & Palmer, 2018). Consistently, empirical research identifies teachers’ limited assessment literacy—understood as teachers’ knowledge and skills related to the design, implementation, interpretation, and use of assessment information to support student learning—as a central constraint for the learning-effective implementation of FA and PASA (Xu & Brown, 2016). To address these issues, a targeted professional development program (FORMA) was designed. Drawing on research on effective teacher professional development, key design features were identified through a review of the teacher education literature and integrated into the program (Oeksuez et al., 2024). FORMA focused on formative PASA in Grades 4–5 mathematics and was implemented as a short multimodal course consisting of three two-hour sessions over three weeks, combining face-to-face and synchronous online formats with asynchronous activities. Building on this framework, the study examines (1) to what extent participation in the FORMA program enhances teachers’ assessment literacy regarding the design, purposes, and perceived benefits of formative PASA, and whether this effect is influenced by teachers’ epistemological beliefs; and (2) to what extent participation in the program enhances the instructional quality of formative peer- and self-assessment in teachers’ mathematics classrooms and whether this relationship is mediated by teachers’ assessment literacy. Methodology, Methods, Research Instruments or Sources Used The study employed a randomized pretest–posttest design with a wait-list control group to examine the effects of a professional development intervention on teachers’ assessment literacy and the instructional quality of PASA within FA. Participants were primary public school teachers teaching mathematics in Grades 4 and 5 in the German-speaking part of Switzerland. Depending on the research question and data availability, sample sizes ranged from 89 to 102 teachers. Following the pretest, participants were randomly assigned to a peer-assessment intervention group, a self-assessment intervention group, or a wait-list control group. Data were collected at pretest and posttest. Teachers’ assessment literacy was assessed using text-based and video-based vignettes with open-ended responses capturing declarative knowledge and situation-specific pedagogical reasoning in FA contexts. Epistemological beliefs were measured using standardized Likert-scale items. Open-ended responses were analyzed using structured qualitative content analysis. Main categories were defined deductively based on FA and PASA theory and refined inductively through data analysis. Response quality was rated on a three-point scale (1 = low, 2 = moderate, 3 = high) according to predefined criteria. Intercoder reliability was high (Cohen’s κ = .82–.90), and interrater reliability for quantitative ratings was excellent (ICC = .86–.91). All questionnaire-based measures were administered jointly within a single survey at each measurement point. To assess the instructional quality of PASA, 96 video-recorded mathematics lessons were analyzed. PASA sequences were identified by coding all lesson segments in which PASA practices occurred and served as the units of analysis. Based on prior research on FA and PASA, a literature-based rating instrument was developed deductively, with indicators refined inductively to capture enacted classroom practices. The instrument demonstrated high interrater reliability (ICC = .90–.92, 95% confidence intervals: .84–.96; all p < .001) and good internal consistency (Cronbach’s α = .86). Intervention effects on teachers’ assessment literacy were examined using repeated-measures ANOVA and ANCOVA. Group differences in the instructional quality of PASA were analyzed using one-way ANOVA with planned contrasts. Mediation analyses tested whether gains in teachers’ assessment literacy mediated the relationship between participation in the professional development program and enacted PASA quality within FA. Conclusions, Expected Outcomes or Findings The intervention led to significant gains in teachers’ assessment literacy (AL), which were limited to declarative knowledge. Teachers in both intervention groups achieved significantly higher scores on the text-based vignette than those in the wait-list control group, with moderate effect sizes. In contrast, no significant improvements were observed for the video-based vignette, indicating that situation-specific assessment skills did not measurably improve. Planned contrasts showed that both intervention groups outperformed the control group, while no significant differences emerged between the peer- and self-assessment conditions. Gains in assessment literacy were positively associated with increases in constructivist epistemological beliefs. Regarding classroom practice, teachers who participated in the intervention implemented formative PASA with significantly higher instructional quality than teachers in the control group. These differences were evident across several central dimensions of PASA enactment. Mediation analyses revealed no significant indirect effects via teachers’ assessment literacy, indicating that improvements in instructional quality were not mediated by the measured knowledge gains. Overall, the findings indicate that even short-term professional development can strengthen teachers’ conceptual understanding of formative PASA and support its implementation in classroom practice, particularly when curriculum-embedded, practice-oriented materials are provided. At the same time, it remains unclear to what extent these effects are sustained over time and transferable to other subject domains, underscoring the need for longitudinal and cross-curricular research. References Andersson, C., & Palm, T. (2018). Reasons for teachers’ successful development of a formative assessment practice through professional development. Studies in Educational Evaluation, 58, 199–206. https://doi.org/10.1080/0969594X.2018.1430685 Black, P., & Wiliam, D. (2018). Classroom assessment and pedagogy. Assessment in Education: Principles, Policy & Practice, 25(6), 551–575. https://doi.org/10.1080/0969594X.2018.1441807 Buholzer, A., B, M., Zulliger von Mühlenen, S., Torchetti, L., Ruelmann, M., Häfliger, A., & Lötscher, H. (2020). Formatives Assessment im alltäglichen Mathematikunterricht von Primarlehrpersonen: Häufigkeit, Dauer und Qualität. Unterrichtswissenschaft. Advance online publication. https://doi.org/10.5281/zenodo.4305566 Cauley, K. M., & McMillan, J. H. (2010). Formative assessment techniques to support student motivation and achievement. The Clearing House: A Journal of Educational Strategies, Issues and Ideas, 83(1), 1–6. https://doi.org/10.1080/00098650903267784 Gotwals, A. W., Philhower, J., Cisterna, D., & Bennett, S. (2015). Using video to examine formative assessment practices as measures of expertise for mathematics and science teachers. International Journal of Science and Mathematics Education, 13(2), 405–423. https://doi.org/10.1007/s10763-015-9623-8 Granberg, C., Palm, T., & Palmberg, B. (2021). A case study of a formative assessment practice and the effects on students’ self-regulated learning. Studies in Educational Evaluation, 68, Article 100955. Maskos, K., Schulz, A., Oeksuez, S. S., et al. (2025). Formative assessment in mathematics education: A systematic review. ZDM Mathematics Education, 57, 679–693. https://doi.org/10.1007/s11858-025-01696-x Narciss, S., Hammer, E., Damnik, G., Kisielski, K., & Körndle, H. (2020). Promoting prospective teacher competencies for designing, implementing, evaluating, and adapting interactive formative feedback strategies. Psychology Learning & Teaching, 20(2), 261–278. https://doi.org/10.1177/1475725720971887 Oeksuez, S. S., Buholzer, A., Maskos, K., & Schulz, A. (2024). Theoriebasierte Entwicklung einer Weiterbildung zu fachspezifischem formativem Peer- und Self-Assessment. In C. Schreiner, G. Schauer, & C. Kraler (Hrsg.), Pädagogische Diagnostik und Lehrer:innenbildung: Bildungswissenschaftliche und fachdidaktische Perspektiven, 107–118. Julius Klinkhardt. Rakoczy, K., Pinger, P., Hochweber, J., Klieme, E., Schütze, B., & Besser, M. (2019). Formative assessment in mathematics: Mediated by feedback’s perceived usefulness and students’ self-efficacy. Learning and Instruction, 60, 154–165. https://doi.org/10.1016/j.learninstruc.2018.01.004 Wanner, T., & Palmer, E. (2018). Formative self- and peer assessment for improved student learning: The crucial factors of design, teacher participation and feedback. Assessment & Evaluation in Higher Education, 43(7), 1032–1047. https://doi.org/10.1080/02602938.2018.1427698 Xu, Y., & Brown, G. T. L. (2016). Teacher assessment literacy in practice: A reconceptualization. Teaching and Teacher Education, 58, 149–162. https://doi.org/10.1016/j.tate.2016.05.010 Yan, Z., Chiu, M. M., & Cheng, E. C. K. (2022). Predicting teachers’ formative assessment practices: Teacher personal and contextual factors. Teaching and Teacher Education, 114. https://doi.org/10.1016/j.tate.2022.103718 09. Assessment, Evaluation, Testing and Measurement
Paper From Foundational Mechanics to Applied Physics: Topic-Sensitive Assessment of Pharmacy Students’ Learning Başkent University, Turkey (Türkiye) Presenting Author:Physics is a compulsory component of many undergraduate pharmacy and health-science programmes, yet assessment practices in these courses often prioritise abstract problem formats that are weakly aligned with students’ professional contexts. From an assessment perspective, this raises a central validity concern: whether observed achievement reflects students’ understanding of physics concepts or primarily reflects their ability to respond to decontextualised task formats. Contemporary validity theory emphasises that the meaning of assessment outcomes depends on the plausibility of the interpretations and uses made from test scores, given the specific tasks, contexts, and intended purposes of assessment (Kane, 2016). This study examines topic-specific assessment outcomes in an undergraduate physics course for pharmacy students, focusing on how performance varies across assessments with different contextual characteristics. The analysis contrasts a midterm examination centred on foundational mechanics with a final examination assessing applied physics topics relevant to pharmaceutical contexts. Rather than treating physics achievement as a single, generalisable construct, the study adopts a context-based assessment perspective, in which performance is understood as situated within specific task contexts and representational demands. Context-based assessment is understood as creating opportunities for resonance between classroom tasks and societal or professional fields, thereby shaping how students mobilise disciplinary knowledge in assessment situations (Bellocchi et al., 2016). The study addresses the following research questions:
The primary objective of the study is to contribute to the interpretation and validity of assessment outcomes in interdisciplinary higher education. In line with argument-based approaches to validity, assessment results are treated as evidence supporting particular interpretive claims about student achievement, rather than as direct indicators of stable or general ability (Kane, 2016). By retaining topic-level scores rather than relying on aggregated exam totals, the study aims to provide a more nuanced account of what physics assessments capture in service-course contexts. Authentic assessment in higher education has been characterised by tasks that prioritise conceptual understanding, problem solving, and disciplinary relevance over procedural recall, particularly in STEM fields (Schultz et al., 2022). Empirical work at the university level also indicates that context-based testing can reveal aspects of students’ conceptual understanding that remain obscured in conventional assessments, especially for non-science majors. Within this framework, different physics topics are treated as distinct assessment domains associated with varying symbolic, procedural, and contextual demands. Consequently, variation in student performance across topics is interpreted as meaningful assessment information rather than as measurement error. The international relevance of the study lies in its focus on pharmacy education, a domain in which physics plays a foundational yet often under-recognised role. Recent work highlights the importance of aligning physics instruction and assessment with pharmaceutical applications to support systems thinking and professional relevance in pharmacy curricula (Dhina et al., 2024; Sakallı, 2024). By empirically examining how assessment context shapes observable achievement within a single course, this study contributes to broader discussions on assessment validity, contextual alignment, and fairness in interdisciplinary higher-education programmes. Methodology, Methods, Research Instruments or Sources Used The study employs a quantitative, observational research design based on assessment data collected from a compulsory undergraduate physics course for first-year pharmacy students at a private university. The course serves as a service module within the pharmacy curriculum and includes both foundational and applied physics content. The study involved 67 first-year undergraduate pharmacy students. Due to missing final examination data, one participant was excluded, yielding a final sample of 66 participants. Two formal assessment instruments were analysed: 1. Midterm Examination (Foundational Assessment Context): The midterm assessed core mechanics topics, including SI units, measurement, forces, one- and two-dimensional motion, torque–equilibrium, and rotational motion. Items were predominantly quantitative and required symbolic manipulation, vector reasoning, and procedural problem solving. These topics reflect traditional physics assessment practices commonly used in early undergraduate instruction. 2. Final Examination (Applied Assessment Context): The final exam assessed applied physics topics with clear relevance to pharmaceutical and health-science contexts, including static electricity, the Bernoulli principle, momentum–impulse, resistors, and capacitors. Tasks required students to apply physics concepts within contextualised scenarios and to integrate conceptual reasoning with quantitative analysis. Both examinations consisted of open-ended problem-solving items scored using standardised analytic rubrics. Rather than relying on total exam scores, topic-level scores were retained for analysis to preserve domain specificity and to support context-sensitive interpretation of achievement. Data analysis proceeded in three stages. First, descriptive statistics were computed to examine score distributions, central tendencies, and variability across topics and assessment contexts. Second, correlation analyses were conducted to explore relationships between foundational and applied topic scores, with particular attention to whether performance patterns were consistent across contexts. Third, an exploratory profile analysis was performed to identify groups of students exhibiting similar cross-topic achievement patterns, thereby supporting interpretation at the individual level. Given the sample size, the analysis focused exclusively on observed variables and avoided latent-variable modelling. This decision aligns with the study’s emphasis on interpretive clarity and assessment validity rather than generalisation or causal inference. All data were collected as part of regular course assessment. Ethical approval was not required under institutional guidelines, as the study involved secondary analysis of anonymised assessment data and posed no risk to participants. Conclusions, Expected Outcomes or Findings The study is expected to demonstrate that pharmacy students’ physics achievement is highly sensitive to assessment context and topic domain, rather than being adequately represented by aggregate exam scores. Specifically, it is anticipated that student performance will differ systematically between assessments focused on foundational mechanics and those assessing applied, discipline-relevant physics topics. Preliminary patterns suggest that achievement in foundational mechanics will show only moderate alignment with performance in applied physics topics. A substantial proportion of students are expected to exhibit weaker performance in abstract, decontextualised mechanics tasks while demonstrating stronger achievement in applied contexts that align with pharmaceutical applications. This finding would support the interpretation that conventional foundational assessments may underestimate competencies that become visible when physics knowledge is assessed through context-based tasks. The analysis is also expected to identify distinct assessment profiles at the individual level. These profiles are likely to include: (1) students demonstrating consistently high performance across both foundational and applied topics; (2) students showing limited achievement in foundational mechanics but strong performance in applied contexts; and (3) students exhibiting persistent difficulties across both assessment contexts. The second profile is anticipated to be particularly prevalent, highlighting the interpretive limitations of relying solely on early, mechanics-heavy assessments in service courses. From an assessment perspective, these findings are expected to contribute to discussions on validity and fairness in interdisciplinary higher education. By showing that topic-level and context-sensitive analyses yield richer information about student achievement than total scores alone, the study underscores the importance of aligning assessment interpretation with task context and intended use. Overall, the study is expected to provide empirical support for context-aware interpretations of assessment outcomes, informing assessment design and evaluation practices in pharmacy and other health-science programmes across European higher-education contexts. References Bellocchi, A., King, D. T., & Ritchie, S. M. (2016). Context-based assessment: creating opportunities for resonance between classroom fields and societal fields. International Journal of Science Education, 38(8), 1304–1342. https://doi.org/10.1080/09500693.2016.1189107 Dhina, M. A., Kaniawati, I., Abdullah, A. G., & Hasanah, L. (2024). Physics in pharmacy education: A socio-technical inquiry into advancing systems thinking and generic science. Journal of Engineering Science and Technology, 19(4), 48–56. Kane, M. T. (2016). Explicating validity. Assessment in Education: Principles, Policy & Practice, 23(2), 198–211. https://doi.org/10.1080/0969594X.2015.1060192 Sakallı, İ. (2024). Integral role of physics in advancing pharmacy education and research. EMU Journal of Pharmaceutical Sciences, 7(3), 122–132. https://doi.org/10.54994/emujpharmsci.1591115 Schultz, M., Young, K., Gunning, T. K., & Harvey, M. L. (2022). Defining and measuring authentic assessment: A case study in the context of tertiary science. Assessment & Evaluation in Higher Education, 47(1), 77–94. https://doi.org/10.1080/02602938.2021.1887811 09. Assessment, Evaluation, Testing and Measurement
Paper Supporting Materials in Statistics Exams: Effects on Achievement and Student Perceptions The Academic College of Tel Aviv-Yaffo, Israel Presenting Author:The use of supporting materials in examinations remains a contested issue in educational assessment, particularly in quantitatively demanding disciplines such as statistics. While closed-book exams are traditionally associated with academic rigor, alternative formats including formula sheets and open-book exams are increasingly adopted to reduce cognitive load and support conceptual understanding. However, empirical evidence comparing different types of support materials within the same assessment context remains limited. This study investigates how different examination formats influence student achievement and assessment experience in an undergraduate introductory statistics course. Specifically, it compares four conditions: closed-book exams, instructor-prepared formula sheets, student-prepared formula sheets, and fully open-book exams. The research addresses two main questions: (1) Do different types of supporting materials affect students’ exam performance and pass rates? and (2) How do these formats influence students’ perceptions of usefulness and preferences regarding assessment design? Grounded in theories of cognitive load and assessment validity, the study conceptualizes supporting materials as tools that may shift assessment focus from memorization to application and reasoning. By examining both performance outcomes and student perceptions, the research contributes to ongoing debates regarding fairness, rigor, and pedagogical effectiveness in assessment practices. The international relevance of this work lies in its implications for higher education instructors seeking evidence-based flexibility in exam design across diverse educational contexts. Methodology, Methods, Research Instruments or Sources Used A randomized between-subjects experimental design was employed with 191 undergraduate students enrolled in an introductory statistics course. Participants were randomly assigned to one of four examination conditions: closed-book, instructor-prepared formula sheet, student-prepared formula sheet, or fully open-book. All students completed the same midterm examination under standardized conditions. Quantitative performance measures included exam scores and pass/fail outcomes. In addition, students completed post-exam survey items assessing perceived usefulness of the exam format and preferences for future assessments. Statistical analyses included descriptive statistics, one-way analysis of variance for exam scores, non-parametric comparisons where appropriate, chi-square tests for pass rates, and logistic regression models examining the relationship between assessment format and likelihood of passing. Perception measures were analyzed using Kruskal–Wallis tests and post hoc comparisons. Ethical approval was obtained in accordance with institutional guidelines, and participation in perception surveys was voluntary and anonymized. Conclusions, Expected Outcomes or Findings Results indicated no statistically significant differences in exam scores or pass rates among the four specific assessment formats. However, when the three supported conditions were combined and compared with the closed-book condition, a modest performance advantage emerged for exams allowing supporting materials. In contrast, substantial differences were observed in students’ perceptions. Participants in all supported formats reported significantly higher perceived usefulness of the assessment and expressed strong preference for the format they experienced. These perception effects were consistent across all types of supporting materials. The findings suggest that while the specific form of support material has limited impact on achievement, access to any form of support plays a central role in shaping students’ assessment experience. This highlights the importance of considering affective and motivational dimensions of assessment alongside performance outcomes. From an assessment design perspective, the results provide empirical evidence that instructors can adopt flexible examination formats without compromising academic standards, while simultaneously enhancing student engagement and confidence. References Sanborn, A. N., & Thuente, D. J. (2012). Learning and memory from open-book tests. Memory & Cognition, 40(4), 509–519. Song, Y., & Thuente, D. J. (2015). The impact of cheat sheets on exam performance. Educational Psychology Review, 27, 1–20. Sweller, J. (1988). Cognitive load during problem solving. Cognitive Science, 12, 257–285. | ||
