Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
Please note that all times are shown in the time zone of the conference. The current conference time is: 19th Aug 2026, 20:15:34 EET
|
Daily Overview |
| Session | ||
09 SES 05 B: Theoretical and Policy Perspectives on Assessment Reform
Paper Session | ||
| Presentations | ||
09. Assessment, Evaluation, Testing and Measurement
Paper Measurement or Judgment? Uncovering the Hidden Paradigms Behind Assessment Reform 1: University of Jyväskylä, Finland; 2: Deakin University, Australia; 3: Kristianstad University, Sweden Presenting Author:Across education systems, pressures to standardise assessment collide with traditions of teacher autonomy. This tension raises a crucial question: What counts as valid assessment? The fundamental nature of assessment is often ignored in policy debates and public discourse. To address this gap, Jönsson (2020) has proposed a distinction between two approaches: assessments based on qualitative (or evaluative) judgments made by assessors and those grounded in quantitative measurement principles informed by psychometrics. In this study, we refer to these paradigms as judgment and measurement. Psychometrics traces its origins to the latter half of the 19th century, with major advances occurring during the 20th century in relation to the measurement of human intelligence. From this foundation, numerous branches have emerged, including tests of scholastic aptitude, personality, and academic achievement (e.g., Black, 1998; Urbina, 2004). Qualitative judgments (Sadler, 1989), also called evaluative judgments (Boud et al., 2018), are used to appreciate the quality of student performance. According to Sadler (1989), a qualitative judgment is made by a knowledgeable expert and is not reducible to a formula that nonexperts can apply. The most notable difference between evaluative judgment and a psychometric perspective lies in their focus. In psychometrics, assessment is treated as an inferential process, targeting an invisible theoretical construct such as knowledge. Consequently, psychometric approaches depend on indirect indicators, aggregating data from multiple items. In contrast, evaluative judgment centers on the quality of performance itself. Assessing an essay or report is a direct act, requiring no inference about hidden traits. Building on this conceptual distinction, this study examines how these paradigms framed summative final assessment and positioned teachers in the context of a nationwide Finnish assessment reform. The reform introduced detailed national assessment criteria for final grades at the end of lower-secondary education, aiming to make teachers’ grading more consistent while preserving classroom-based assessment practices. A clear understanding of final assessment is essential because of its profound impact on students’ futures and the functioning of society. In this context, teachers are users of public power and may need to justify their grading decisions. Clarity in grading principles is critical: Should professional judgment count as valid evidence, or should grades rely solely on external measures? Do we want teachers to act as assessment machines? Or do we want to externalise assessment to tests, viewing teachers as technicians without power over assessment? We analyzed two complementary datasets: 56 official and media documents saping public discourse during the reform, and 66 semi-structured interviews with lower-secondary subject teachers across six disciplines. Our analysis of the public and official documents showed that the measurement and judgment paradigms were intertwined in discourse, and so tightly that it was difficult to disentangle them. Moreover, teachers were positioned in somewhat confusing ways both 1) as parts of the assessment machine (conductors of measurement) and 2) as assessment machines (professionals using evaluative judgment). In the interview data, we identified three distinct ways in which teachers referred to assessment and their positionings within it: 1) as designers of the assessment machine, who constructed instruments to ensure reliability (referring to measurement), 2) as assessment machines, who applied holistic, criteria-based judgments that resisted reduction to mere metrics (referring to judgment); and 3) as calibrators who adjusted quantitative indicators intuitively with the criteria (referring to intuitive judgment). Our findings indicate that these paradigms of assessment are intertwined in societal discourse, simultaneously positioning teachers as agentic professionals and sources of error – an inherently contradictory stance. This underscores the need for clearer conceptualisations of assessment to inform both policy decisions and practical implementation, clarifying whether teachers are expected to act as objective technicians or professional interpreters. Methodology, Methods, Research Instruments or Sources Used We conducted a qualitative, theory-informed analysis combining (1) documentary and media sources and (2) teacher interviews. The document data (n = 56) included official guidance and reports from the Finnish National Agency for Education (FNAE) and the Ministry of Education and Culture (MEC), alongside prominent news coverage from Helsingin Sanomat (the most widely published newspaper in Finland) and Yleisradio (Finland’s national public service broadcaster) during 2017–2024. To identify relevant documents, we used the search term “päättöarv*” (final assess*) across both outlets for the years 2017–2024. To understand how teachers themselves related to the measurement and judgment paradigms, we analysed teacher interviews that were collected as part of the nationwide research project funded by the MEC. The interview sample comprised 66 lower-secondary subject teachers from six disciplines (history, mathematics, physics, mother tongue and literature, English as a foreign language, Finnish/Swedish as a second language). Interviews were semi-structured, discipline-informed, and averaged 53 minutes (range: 23–87 minutes), covering experiences of the reform, criteria use, grading practices, and perceived impacts. Participation was voluntary with informed consent. We operationalized two paradigms of assessment—psychometric measurement and evaluative judgment—drawing on Jönsson’s (2020) conceptual paper. Indicators of a psychometric paradigm included quantitative scoring, summarizing scores from several items, inferences about latent traits, primary focus on reliability, and low teacher autonomy. Indicators of and evaluative judgment paradigm included assessment of performance quality, holistic use of several criteria simultaneously, direct (non-inferential) assessment, primary focus on validity, and high teacher autonomy. First, given the large volume of data, we began by extracting parts that addressed assessment as a psychometric practice or a process of evaluative judgment. Next, based on these excerpts, we analyzed how and by whom these two paradigms were produced and how they were potentially contested across the datasets. Using the abovementioned indicators, we systematically mapped how final assessment was conceptualised throughout the reform. Finally, we examine how teachers were positioned in relation to assessment (Harré et al., 2009), focusing on how ideal, normal and avoidable teacher positionings were articulated by different actors and documents, as well as by teachers themselves. Conclusions, Expected Outcomes or Findings Our study shows that the discourse around assessment reform entwined the two distinct conceptualizations of assessment: psychometric measurement and evaluative judgment. Public and official communications simultaneously affirmed teacher professionalism and questioned teacher reliability. This double discourse framed teachers as both expert interpreters and potential sources of error, creating an inherently paradoxical positioning. Similarily, teachers’ accounts reflected various conceptualizations of assessment and assessors. These illustrate adaptive professionalism but also reveal conceptual tensions: when policy discourse merges psychometric measurement with evaluative judgment, teachers receive ambiguous signals about what constitutes valid evidence and how final grades should be determined. We argue that conceptual clarity is essential for both policy and practice. Assessment systems should articulate when a psychometric orientation is warranted and when evaluative judgment is central. Professional development practices should be designed accordingly. If the differences between the paradigms are not understood, and if measurement is framed as a superior assessment paradigm or used as a benchmark for teachers’ judgments, this can contribute to mistrust in teacher professionalism (see Novak & Carlbaum, 2017). The conflicting expectations placed on teachers – between acting as objective technicians and professional interpreters – echoes long-standing global tensions in educational assessment (e.g., Broadfoot & Black, 2004; Klenowski, 2011), particularly in how assessment is understood publicly (Gardner, 2017). Finland offers a compelling case due to its strong tradition of teacher autonomy. That teachers feel both empowered and scrutinized in this context suggests that external pressures to standardize and quantify assessment (often introduced by global actors such as the OECD) can reshape professional identities even within high-autonomy systems. Ultimately, assessment does not merely measure learning; it organizes professional work and identities. The Finnish case suggests that sustaining trust in teacher-determined final grades requires naming the paradigms at play and designing infrastructures that help teachers judge well. References Black, P. (1998). Testing: Friend or foe? Theory and practice of assessment and testing. Routledge. Boud, D., Ajjawi, R., Dawson, P., & Tai, J. (2018). Developing evaluative judgement in higher education. Routledge. Broadfoot, P., & Black, P. (2004). Redefining assessment? The first 10 years of assessment in education. Assessment in Education: Principles, Policy & Practice, 11(1), 7–26. https://doi.org/10.1080/0969594042000208976 Gardner, J. (Ed.). (2017). The public understanding of assessment. Routledge. Harré, R., Moghaddam, F. M., Cairnie, T. P., Rothbart, D., & Sabat, S. R. (2009). Recent advances in positioning theory. Theory & Psychology, 19(1), 5–31. https://doi.org/10.1177/095935430810 Jönsson, A. (2020). Definitions of formative assessment need to make a distinction between a psychometric understanding of assessment and “evaluative judgment.” Frontiers in Education, 5. https://doi.org/10.3389/feduc.2020.00002 Klenowski, V. (2011). Assessment for learning in the accountability era: Queensland, Australia. Studies in Educational Evaluation, 37(1), 78–83. https://doi.org/10.1016/j.stueduc.2011.03.003 Novak, J., & Carlbaum, S. (2017). Juridification of examination systems: Extending state level authority over teacher assessments through regrading of national tests. Journal of Education Policy, 32(5), 673–693. https://doi.org/10.1080/02680939.2017.1318454 Sadler, D. R. (1989). Formative assessment and the design of instructional systems. Urbina, S. (2004). Essentials of psychological testing. John Wiley and Sons. 09. Assessment, Evaluation, Testing and Measurement
Paper A Missed Window of Opportunity for Assessment Reform in Lithuania: Applying the Multiple Streams Framework Vytautas Magnus University, Education Academy Presenting Author:In Lithuania, a new competency-based curriculum has recently been introduced alongside a two-stage external state examination system. The intention was to replace high-stakes examinations, predominantly composed of closed- or short-answer items, with competency-based assessments that give students an opportunity to demonstrate their unique skills and performance. However, the newly developed examination system did not deliver the expected results. The number of examinations increased when examination parts were transferred to a digital format, the test design remained unchanged, and no performance assessment was introduced (Raudiene et al., 2025). After the first year of implementing the new exam, marked by a high number of issues related to poor item quality, difficulty-level mismatches, and technical implementation, which raised public dissatisfaction and anger, we decided to analyze how the decision-making process was carried out and how far the current assessment system is from the competency-based assessment. Focusing on competencies as learning outcomes assumes that checking whether students recall information is no longer the highest priority, while assessing what students can do with knowledge is (Darling-Hammond & Adamson, 2010; Elmer et al., 2019; Black & Wiliam, 1998). Assessing how students apply knowledge in new situations goes beyond traditional paper-and-pencil tests or their digital alternatives and involves direct interaction between the assessor and the assessed, which may not consistently yield quantitative feedback. This does not satisfy the bureaucratic apparatus that uses exam scores, indicators, and rankings for governance and accountability at all levels, from students and teachers to municipalities and states (Ball, 2003). Little is questioned about what this culture of measurement brings to education apart from competition, stratification, the distortion of basic pedagogical principles, and distrust in education (Biesta, 2010; Koretz, 2017; Au, 2009). Because of the lack of clear rules and a diverse array of actors, the development of education policy is complex (Shaw, 2019). Ball (1998) described it as a process of bricolage, borrowing and copying bits of ideas from elsewhere, interpreting theories and research, and reusing them as though they might work. Due to the high stakes associated with school examinations, university enrollment, school rankings, and follow-up consequences, assessment policy-making is considered extremely complex and multilayered, thus requiring a well-established theoretical approach. The Multiple Streams Framework (MSF) (Kingdon, 2014) offers a lens for understanding how certain proposals reach the political agenda and why some issues receive political attention at certain times while others do not. The MSF is a tool for analyzing nonlinear and non-rational political processes (Shaw, 2019; Stout & Stevens, 2000), including policy agenda-setting, identifying the necessary conditions for translating political agendas into concrete decisions, determining the timing of reform, and diagnosing policy failures. Kingdon (2014) suggests that policy development comprises three streams:
These streams develop independently, but at critical moments all three conflate, opening a window of opportunity to introduce a new policy. Policy-making is enhanced by policy entrepreneurs who invest resources, such as expertise, political influence, and determination, to bring together all three streams to achieve policy change. Methodology, Methods, Research Instruments or Sources Used For this study, we adopted a qualitative research design to examine how MSF could explain the process of assessment policy-making and the formation of assessment practices in Lithuania. Recognizing that educational assessment affects a large number of people from diverse social groups, including students, teachers, parents, headteachers, academics, municipal and national-level education managers, politicians, government officials, publishers, and others, we sought to reach representatives of all these groups and conducted 24 in-depth interviews. Among our participants were the highest-level government officials directly responsible for introducing new assessment policies, as well as other representatives who are publicly known for their positions on assessment policy. These interviews were conducted using the MS Teams platform between August and October 2025. Following MSF, we asked all our participants the same three questions: (1) how they perceived the critical problem in educational assessment; (2) what alternative solutions were on the table; and (3) what the political context was for implementing new policies. We asked for their reflections from their current positions, as some had already left their previous posts. To analyze the collected data, we used thematic analysis, which is known for reflecting reality and revealing its surface (Braun & Clark, 2006), making it a perfect match for Kingdon’s MSF. When conducting thematic analysis of the collected data, we followed a six-stage approach: familiarization with the data through careful reading and preparation for analysis, initial code generation that involves identifying segments in the raw data that can be assessed in a meaningful way (Boyatzis, 1998), searching for themes by grouping codes into potential themes, reviewing themes by revising them, and then defining themes by identifying the “essence of what each theme is about” (Braun & Clark, 2006, p. 22). Our interpretative work through the first five stages helped identify the recurring analytical patterns and start the final sixth stage of report writing. The MSF provided a structure for our narrative, as the identified themes were organized around three MSF streams - problems, policy, and politics. Alongside the policy entrepreneurs were named, and our evaluation focused on whether and how an open policy window was used Conclusions, Expected Outcomes or Findings Our analysis identified critical turns that shaped assessment policymaking: Within the problem strand, the common concerns focused on: (1) misalignment between the expected learning outcomes and the examination system, and on failed promises to introduce competency-based assessment; (2) equity and fairness issues were repeatedly raised, especially regarding the assessment experiences of disadvantaged student groups; (3) the quality and consistency of exam design were discussed, with reference to a lack of expertise in test development. The policy strand revealed possible alternatives: (1) growing support for systemic change to university admission, proposing more pathways to the university and not limiting acceptance to examination results alone; (2) improving the effectiveness of the examination system by introducing assessment innovations such as comparative judgment, assessment moderation, performance-based and adaptive assessments, etc. The politics strand revealed that: (1) not only is strong political leadership and engagement by top government officials crucial, but that without sufficient parliamentary support, things might get complicated; (2) certain bureaucratic procedures and a lack of expertise among the permanent staff in the Ministries and Agencies might become weak spots leading to failure; (3) different stakeholder groups emphasized different aspects of the problem, requiring distinct solutions; for example, students are mostly concerned with how well they pass exams and university entrance; headteachers are disappointed with school rankings; administrative staff are preoccupied with the technical implementation; while government officials worry about how their decisions will affect future election results. This reminded us of a battle where all the fighters had different ideas of what their victory would be like. Finally, we conclude that although curriculum reform created a policy window for political change to align learning goals with assessment approaches, the lack of a shared view among policy entrepreneurs on how to achieve the desired goal left the window unopened. References Au, W. (2009). Unequal by design: High-stakes testing and the standardization of inequality (1st ed.). Routledge. Ball, S. J. (1998). Big policies/small world: An introduction to international perspectives in education policy. Comparative Education, 34(2), 119–129. https://doi.org/10.1080/03050069828225 Ball, S. J. (2003). The teacher’s soul and the terrors of performativity. Journal of Education Policy, 18(2), 215–228. https://doi.org/10.1080/0268093022000043065 Biesta, G. (2017). Education, Measurement and the Professions: Reclaiming a space for democratic professionality in education. Educational Philosophy and Theory, 49(4), 315–330. https://doi.org/10.1080/00131857.2015.1048665 Black, P., & Wiliam, D. (1998). Assessment and classroom learning. Assessment in Education: Principles, Policy & Practice, 5(1), 7–74. https://doi.org/10.1080/0969595980050102 Boyatzis, R. E. (1998). Transforming qualitative information: Thematic analysis and code development. Thousand Oaks, CA: Sage. Braun, V., & Clarke, V. (2006). Using thematic analysis in psychology. Qualitative Research in Psychology, 3(2), 77–101. https://doi.org/10.1191/1478088706qp063oa Darling-Hammond, L. & Adamson, F. (2010). Beyond basic skills: The role of performance assessment in achieving 21st century standards of learning. Stanford, CA: Stanford University, Stanford Center for Opportunity Policy in Education. Emler, T. E., Zhao, Y., Deng, J., Yin, D., & Wang, Y. (2019). Side effects of large-scale assessments in education. ECNU Review of Education, 2(3), 279–296. https://doi.org/10.1177/2096531119878964 Kingdon, J. W. (2014). Agendas, alternatives, and public policies (2nd ed.). Pearson Education. Koretz, D. (2017). The testing charade: Pretending to make schools better. University of Chicago Press. Raudienė, I., Ališauskienė, S., & Mažeikienė, N. (2025). Uncovering the meanings of educational assessment: Realities at the outset of inclusion reform in Lithuanian schools. European Journal of Special Needs Education. https://doi.org/10.1080/08856257.2025.2580903 Shaw, R. D. (2019). Examining arts education policy development through policy frameworks. Arts Education Policy Review, 120(4), 185–197. https://doi.org/10.1080/10632913.2018.1468840 Stout, K. E., & Stevens, B. (2000). The case of the failed diversity rule: A multiple streams analysis. Educational Evaluation and Policy Analysis, 22(4), 341–355. 09. Assessment, Evaluation, Testing and Measurement
Paper Negotiating Professionalism when Assessment Becomes High-stakes Stockholm university, Sweden Presenting Author:Assessment has become an increasingly central aspect of teachers’ professional work, reshaping how teaching, learning, and evaluation are understood in everyday school practice. Scholars in education around the world suggest how teachers have to respond to these accountability claims in order to show effectiveness and quality, ultimately making school results the focus of education (Biesta, 2009; Irisdotter Aldenmyr et al., 2025). This has affected the teaching profession in several regards, changing core aspects of the institutional structures for teaching and even what it means to be a teacher (Ball, 2003). This paper explores how teachers in Sweden navigate professionalism in relation to assessment in a school system characterized by a focus on school results. Drawing on focus group interviews with teachers across primary, lower secondary, and upper secondary education, the study aims at exploring how teachers in Sweden understand and enact professionalism in assessment, with particular attention to how their professional judgement is negotiated within institutional conditions shaping teaching and learning. When the quality of schools, and even teachers, is defined in terms of measurable educational outcomes, there is a risk of teaching focusing on testing and grading, or what has been labelled “teaching to the test”, which, in turn, leads to tests gaining an ever-increasing role (Vogt, 2022), constructing a testing culture in schools (Birenbaum, 2016). Additionally, a school characterized by performance activities and excessive summative assessments is problematic as it often causes teacher stress (Lundahl, et al., 2022; Saeki et al., 2018; von der Embse et al., 2015) and put students learning and well-being at risk (Brolin Låftman et al., 2013; Hirsh, 2020; Luthar et al., 2020). The study is grounded in sociological theories of professionalism and professional work, combined with research on teachers’ assessment practices and assessment literacy. Professionalism is understood not merely as the possession of formal knowledge or skills, but as a socially and institutionally situated practice shaped by norms, values, and expectations regarding appropriate professional behaviour (Freidson, 2001; Evetts, 2006, 2013). From this perspective, what counts as “professional” teaching is continuously negotiated within specific organizational and policy contexts. A key distinction is made between professionalism as enacted from within the profession and organizational professionalism which is defined from outside of the profession and based on bureaucratic logics and accountability activities (Evetts, 2013). This distinction is used to analyse how teachers’ assessment practices are influenced by institutional demands while still relying on teachers' professional judgement. Teacher agency is conceptualized as relational and context-dependent, drawing on an ecological understanding of agency (Priestley et al., 2015). Teachers’ responses to assessment practices are thus seen as negotiated rather than uniform. To capture the complexity of assessment work, the analysis is further informed by the concept of assessment literacy (Pastore & Andrade, 2019), which distinguishes between conceptual, praxeological, and socio-emotional dimensions of assessment. This framework enables an examination of how teachers balance formal requirements with pedagogical concerns not forgetting ethical considerations and the pupils’ socio-emotional development. Together, these perspectives frame teacher professionalism in assessment as a dynamic process shaped by institutional conditions and pedagogical values. Methodology, Methods, Research Instruments or Sources Used The study is based on qualitative focus group interviews with 39 teachers from 12 schools located in a large urban area in Sweden. The participants work across different school levels, including primary, lower secondary, and upper secondary education, and all hold formal teaching qualifications. Additionally, most teachers have extensive professional experience. The diversity of school contexts enabled exploration of shared and contrasting understandings of assessment across educational stages as teachers in Sweden don’t’ put grades before grade 6. Focus group interviews were chosen to facilitate collective reflection on assessment as a shared professional practice and to capture how teachers construct and negotiate professional meanings in interaction. The group format also provided a supportive setting for discussing experiences that may be sensitive or contested (Bloor et al., 2001). All interviews were audio-recorded and transcribed. The empirical material was analysed using thematic analysis. The analysis proceeded in three stages: familiarization with the data through repeated reading and listening; identification of recurring and salient themes in teachers’ narratives; and interpretative analysis guided by the study’s aim and theoretical framework. Conclusions, Expected Outcomes or Findings The study shows how teachers’ professionalism in assessment is continuously negotiated within institutional boundaries. Teachers express frustration that teaching is being pushed aside by a growing emphasis on summative assessment, ultimately reducing school performance to the primary purpose of education. At the same time, they comply with these assessment demands, even if with a professional unease (c.f. Moore et al., 2002). I interpret this as a split between being able to exercise one's professionalism and being recognized for it. The risk is that teachers are pushed into performative practicies to assert professionalism when summative assessment (grades and test scores) become not only a goal of education per se, but also become a definition of school (and teacher) quality. Yet, the constant need to prove one's professionalism and adjust practices to meet new accountability requirements, risks generating what Ball (2003) terms ontological insecurity. The study also reveals how teachers’ professionalism in assessment is characterized by tensions between different aspects of assessment rather than alignment. Teachers experience conflicts between pedagogical and ethical commitments, such as supporting learning, motivation, and well-being, and assessment practices that prioritize formal regulations, expectations from “outside” and comparability. These tensions are especially evident between formative assessment intended to support learning and summative requirements with, for the teachers, sometimes diffuse purposes. When school results become the primary sign of school quality, professionalism recognized through visible, documentable actions rather than through pedagogical reasoning and ethical judgement, teachers’ professionalism in assessment is being contested. References Ball, S. J. (2003). The teacher’s soul and the terrors of performativity. Journal of Education Policy, 18(2), 215–228. Biesta, G. (2009). Good education in an age of measurement: on the need to reconnect with the question of purpose in education. Educational Assessment, Evaluation and Accountability, 21, 33–46. Birenbaum, M. (2016). Assessment culture versus testing culture: The impact on assessment for learning. I: Assessment for learning: Meeting the challenge of implementation, 275–292. Springer International Publishing. Bloor, G., Frankland, J., Thomas, M. & Robson, K. (2001). Focus Groups in Social research. Sage. Brolin Låftman, S., Almquist, Y. B. & Östberg, V. (2013). Students’ accounts of school-performance stress: a qualitative analysis of a high-achieving setting in Stockholm, Sweden. Journal of Youth Studies, 16, 932–949. Evetts, S. (2013). Professionalism: Value and ideology. Current Sociology Review, 61(5-6), 778–796. Evetts J (2006) Short note: The sociology of professional groups. Current Sociology 54(1), 133–143. Hirsh, Å. (2020). When assessment is a constant companion: students’ experiences of instruction in an era of intensified assessment focus. Nordic Journal of Studies in Educational Policy, 6, 89–102. Irisdotter Aldenmyr, S., Gradén, M. & Håkansson J. (2025). Teacher leaders’ capacity for Luthar, S., Kumar, N.L & Zillmer, N. (2020). Teachers’ responsibilities for pupils’ mental health: Challenges in high achieving schools. International journal of school & educational psychology, 8(2), 119–130. Lundahl, C., Mickwitz, L. & Skott, P. (2022). Skolutveckling för hållbart lärande – teoretiska och praktiska perspektiv. Studentlitteratur. Moore, A., Edwards, G, Halpin, D. & George, R. (2002). Compliance, Resistance and pragmatism: the (re)construction of schoolteacher identities in a period of intensive educational reform. British Educational Research Journal, 28(4). Pastore, S. & Andrade, H. (2019). Teacher assessment literacy: A three-dimensional model. Teaching and Teacher Education, 84, 128-138. Priestley, M.; Biesta, G. & Robinson, S. (2015). Teacher Agency: An Ecological Approach. Bloomsbury Academic. Saeki, E., Segool, N. & Pendergast, L. (2018). The influence of test-based accountability policies on early elementary teachers: School climate, environmental stress, and teacher stress. Psychology in the Schools 55(4). Vogt, B. (2022). Supportive assessment strategies as curriculum events in a performance-oriented classroom context. European Educational Research Journal, 21(6), 1023-1040. von der Embse, N. P., Kilgus, S. P., Solomon, H. J., Bowler, M., & Curtiss, C. (2015). Initial Development and Factor Structure of the Educator Test Stress Inventory. Journal of Psychoeducational Assessment, 33(3), 223-237. 09. Assessment, Evaluation, Testing and Measurement
Paper Operationalising Performance Level Descriptors for Classroom Assessment: From Generic Scales to Warranted Teacher Judgement University of Sydney, Australia Presenting Author:Performance Level Descriptors (PLDs) are embedded in curriculum frameworks internationally to describe the quality of student achievement (Cizek & Bunch, 2007; Sadler, 1987). These frameworks respond to legitimate concerns about the limitations of numerical scores and provide a common language for describing achievement. However, existing approaches to PLDs face a persistent dilemma: they tend to be either too generic to guide classroom assessment or too holistic to provide diagnostic information (Sadler, 2009; Wyatt-Smith, Klenowski, & Colbert, 2014). Methodology, Methods, Research Instruments or Sources Used This paper reports the theoretical and design phases of a design-based research program (McKenney & Reeves, 2012). It does not report an empirical study of student performance or teacher judgement; rather, it develops a theoretical framework, and demonstrates its application through iterative task design. Problem identification emerged through professional engagement with PLDs in the NSW context, where teachers are expected to assess and report against PLD levels but receive limited guidance on operationalising generic descriptors for specific tasks. While argument-based validity frameworks (Kane, 2013) provide robust tools for evaluating large-scale assessments, and while PLD development literature addresses system-level design, little work addresses how teachers might construct warranted inferences from classroom tasks to PLD placements. This gap exists wherever standards-referenced assessment is implemented. Framework development proceeded through design-based research (McKenney & Reeves, 2012), combining the theoretical synthesis described above with iterative refinement through teacher focus groups across Key Learning Areas (n=52 primary and secondary teachers). Teacher professional knowledge was positioned not as implementation feedback but as a theoretical contribution, with disciplinary expertise revealing where universal cognitive frameworks required translation to respect epistemological differences (Kesidou, 2025). This synthesis produced six components describing lenses through which to view quality in student performance - knowledge integration and use, conceptual understanding, reasoning and metacognition, disciplinary practice, skill fluency, and communication - each grounded in specific research traditions. Critically, we then asked: what would it mean for a PLD placement to be warranted? Drawing on Toulmin's (2003) argument structure, we recognised that generic components alone cannot warrant inference. Components must combine with specific curriculum outcomes to produce an intersection - where a component like 'reasoning and metacognition' meets an outcome like 'compare conservation and sustainability' - that is the actual target of assessment. It is at this intersection that warrants must be articulated. Demonstration involved developing assessment tasks and specifications documenting target inferences. This was iterative: initial designs failed to create conditions for the intended inferences, requiring reconceptualisation of tasks and stimulus materials - evidence of the framework's capacity to reveal problems invisible in generic task design. Next steps involve empirical validation: testing whether teachers can use the specification structure reliably and whether intersection-level descriptors differentiate student performance as predicted. This work is currently underway. Conclusions, Expected Outcomes or Findings This paper addresses a problem common across standards-referenced systems: the gap between generic level PLDs and the specific inferences teachers must make in classroom practice. The analysis yields two primary results: first, theoretically-grounded PLD components as lenses for observing quality in student performance; second, a specification structure for documenting warranted inference at the PLD component - curriculum outcome intersection. Three implications follow. For assessment theory: The component-outcome intersection resolves the generic/holistic dilemma by providing a unit of assessment that is both specific (tied to particular curriculum outcomes) and differentiated (distinguishing components of performance). The framework extends argument-based validity approaches (Kane, 2013) to teacher judgement - not to burden teachers with documentation - but to clarify the inferential structure underlying assessment. This contributes to international dialogue on assessment quality, where tensions between standardisation and professional judgement remain contested (e.g., Black & Wiliam, 1998; Dolin & Evans, 2017; Harlen, 2005). For practice: Teachers need more than generic scales or holistic profiles. This framework offers theoretically grounded PLD components that provide lenses for observing student performance and a specification structure for documenting how these components manifest for particular curriculum outcomes. This creates artefacts that can be shared, critiqued, and refined through collegial moderation - making explicit the reasoning that expert teachers often hold tacitly (Klenowski & Wyatt-Smith, 2014). For policy: If valid application requires outcome-specific operationalisation, curriculum & assessment authorities in standards-referenced systems should demonstrate this process through case studies supporting teacher professional learning. The conference theme calls for dialogue between researchers, practitioners, and decision-makers, and for recognition of diverse knowledge forms. This framework enacts that dialogue: making research-informed assessment reasoning accessible to teachers while drawing on their disciplinary expertise as a theoretical resource. A PLD placement is an inference, made in context, requiring the kind of situated professional knowledge the theme foregrounds. References Barnett, S. M., & Ceci, S. J. (2002). When and where do we apply what we learn? A taxonomy for far transfer. Psychological Bulletin, 128(4), 612–637. Black, P., & Wiliam, D. (1998). Assessment and classroom learning. Assessment in Education: Principles, Policy & Practice, 5(1), 7–74. Cizek, G. J., & Bunch, M. B. (2007). Standard setting: A guide to establishing and evaluating performance standards on tests. Sage. Dolin, J., & Evans, R. (Eds.). (2017). Transforming assessment: Through an interplay between practice, research and policy. Springer. Harlen, W. (2005). Teachers' summative practices and assessment for learning: Tensions and synergies. The Curriculum Journal, 16(2), 207–223. Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1–73. Kesidou, S. (2025, December). Designing for quality across the curriculum: A tiered performance standards framework. Paper presented at the Australian Association for Research in Education (AARE) Annual Conference, Newcastle, Australia. Klenowski, V., & Wyatt-Smith, C. (2014). Assessment for education: Standards, judgement and moderation. Sage. Liu, O. L., Rogat, A., & Bertling, J. P. (2013). A CBAL science model of cognition. ETS Research Report Series. Mayer, R. E. (2011). Applying the science of learning. Pearson. McKenney, S., & Reeves, T. C. (2012). Conducting educational design research. Routledge. Paul, R., & Elder, L. (2019). The miniature guide to critical thinking: Concepts and tools (8th ed.). Foundation for Critical Thinking. Sadler, D. R. (1987). Specifying and promulgating achievement standards. Oxford Review of Education, 13(2), 191–209. Sadler, D. R. (2009). Indeterminacy in the use of preset criteria for assessment and grading. Assessment & Evaluation in Higher Education, 34(2), 159–179. Toulmin, S. E. (2003). The uses of argument (Updated ed.). Cambridge University Press. Wyatt-Smith, Klenowski, V., & Colbert, P. (Eds.), (2014). Designing assessment for quality learning. Springer. | ||
