Conference Agenda
| Session | ||
09 SES 12 B: AI, Trust and Governance in Assessment Practice
Paper Session | ||
| Presentations | ||
09. Assessment, Evaluation, Testing and Measurement
Paper Reliability and Validity of using Generative AI for Qualitative Data Analysis University of Nicosia, Cyprus Presenting Author:The growing popularity of large language models (LLMs) in education has spurred their use in a range of tasks, including creating assessments, analyzing data, and evaluating student responses. However, due to the potential for AI to generate misleading outputs, the European Commission (2024) has proposed guidelines emphasizing responsible usage of generative AI in research, focusing on reliability, honesty, respect, and accountability. These guidelines call for verifying and reproducing information produced by AI, ensuring transparency regarding AI use, recognizing the technology’s limitations, and taking overall responsibility for the outputs. Building on these principles, the present study aimed to demonstrate a process for applying these guidelines in research, centering on evaluating how accurately LLMs reproduce a qualitative data analysis conducted by human researchers and how consistent their outputs are across multiple LLMs. The study posed three core research questions:
Methodology, Methods, Research Instruments or Sources Used The study was carried out in four stages. First, the data were gathered from 66 high-school students through an online questionnaire where the students were asked questions regarding how they believed that AI would be integrated into math and science in schools in the future. Second, the team performed a thematic analysis on the students’ responses, which produced three themes: (1) creation of exercises for practice, (2) interactive learning (including games, experiments, and simulations), and (3) demonstrations of different problem-solving approaches to support students’ understanding. In the third stage, the researchers tested whether various LLMs could recover and categorize the same themes. To account for potential variability across models and across repeated prompting of the same model, the team ran these prompts 10 times on 13 LLMs, including ChatGPT o1 Pro, ChatGPT o3, ChatGPT o3 mini (H), ChatGPT o1, Claude 3.7 Sonnet, ChatGPT o1 mini, ChatGPT 4.5 preview, Gemini 2.o Pro, Gemini 2.0 flash, Deepseek-R1, ChatGPT 4o, Llama 3.1, and Llama 3.3. Finally, the researchers objectively measured each model’s fidelity to the human-established categories by checking whether the LLMs’ categories corresponded to those generated by researchers, and whether the reported number of matched instances aligned with the actual matches. Conclusions, Expected Outcomes or Findings Most LLMs demonstrated strong consistency. ChatGPT o1 Pro, ChatGPT o1, ChatGPT o3, Claude 3.7 Sonnet, ChatGPT o1 mini, ChatGPT 4.5 preview, Gemini 2.0 flash, and ChatGPT 4o perfectly replicated the three categories in all 10 iterations. Some models produced minor divergences: ChatGPT o3 mini (H) struggled slightly with “Demonstrations of multiple problem-solving methods,” reaching 7 out of 10 correct categorizations. Llama 3.1 aligned with the human-defined categories for “Creation of exercises for practice” 8 times out of 10, but only 6 times for the other two themes. Deepseek-R1 and Llama 3.3 completed just three runs, even though prompted for 10, yet each correctly reproduced the categories in all three attempts. Gemini 2.0 Pro presented a separate issue. While it produced 10 sets of 3 categories each (totaling 30 categories), its summary table listed more than 30 instances. Specifically, it recorded 12 instances of “Creation of exercises for practice,” 11 of “Interactive Learning,” and 12 of “Demonstrations,” suggesting redundant categorization or a misunderstanding of the prompt. The average reliability within each LLM equaled 0.95, while Cohen’s Kappa ranged from 0.92-0.96 according to the themes that were analyzed. Further analyses will also be performed to compare the LLM outcomes from a validity perspective. References European Commission (2024). Living guidelines on the responsible use of generative AI in research (ERA Forum Report, March 20, 2024). Brussels: European Commission. 09. Assessment, Evaluation, Testing and Measurement
Paper Confident, Cautious, or Resistant? Educators’ Engagement with AI-Driven Formative Assessment 1: Taylors University Malaysia; 2: University of South Eastern Norway; 3: Coventry University, United Kingdom Presenting Author:Despite strong evidence supporting the value of formative assessment for teaching and learning (Black & Wiliam, 1998), its implementation remains challenging. Studies show that the uptake of formative assessment practices is considerably influence by various factors, and its uptake is often difficult and superficial (Wang et al., 2021). These challenges are commonly described in literature as stemming from both personal and contextual factors (Bennett, 2011). These personal and contextual challenges are compounded by a strong assessment culture and accountability pressures (Asian Development Bank, 2011; Simper et al., 2022). In Asian context, studies suggest that educators’ practices remain predominantly summative, with formative strategies often used superficially or inconsistently due to contextual constraints such as workload, time pressure, and limited professional support (Sulaiman et al., 2020). In recent years, digital technologies and more recently artificial intelligence (AI) have been increasingly positioned as potential means to support formative assessment practices by addressing some of these long-standing challenges, such as providing timely feedback, supporting task design, and managing workload (Hopfenbeck et al., 2023). However, the integration of AI into formative assessment does not occur automatically, nor it is shortcut for learning but these new forms of learning are a catalyst for deeper learning (Chatfield, 2025). Educators’ decisions to adopt, adapt, or resist AI-driven formative assessment tools are shaped by their familiarity with AI, their confidence in using such technologies, and the broader institutional and ethical conditions within which they work. Despite growing interest in AI in higher education, there remains limited empirical evidence on how educators’ AI literacy competencies relate to the adoption of AI-driven formative assessment in practice. This study addresses this gap by investigating the relationship between higher education educators’ self-perceived AI literacy competency and their adoption of AI-driven formative assessment, and by exploring the factors that influence such adoption. Two research questions guide the study: (1) Is there a relationship between educators’ self-perceived AI literacy competency and their adoption of AI-driven formative assessment? and (2) What factors influence higher education educators’ adoption of AI-driven formative assessment tools? The study is theoretically informed by Cultural-Historical Activity Theory (CHAT), which conceptualises learning and professional practice as socially situated and mediated by tools within activity systems (Engeström, 2001). From this perspective, AI is understood as a mediating tool whose adoption is shaped by interactions between subjects (educators), objects (formative assessment practices), rules (institutional policies and norms), community (colleagues and students), and division of labour (support structures and responsibilities). CHAT provides a useful lens for examining how educators’ familiarity with AI develops and how social, institutional, and ethical conditions influence decisions to confidently adopt or cautiously use AI in their formative assessment practices. Methodology, Methods, Research Instruments or Sources Used A mixed-methods research design was employed. Quantitative data were collected through an online survey administered to higher education educators in Malaysia (n = 84). The survey measured self-perceived AI literacy competency across four dimensions: knowledge, awareness, confidence, and usage/adoption. Statistical analyses, including correlation analysis, were conducted to examine relationships between AI literacy dimensions and the adoption of AI-driven formative assessment tools. Qualitative data were collected through semi-structured interviews with a purposive subsample of educators (n = 9). The interviews explored educators’ experiences with predictive and generative AI tools, their formative assessment practices, and perceived enabling and constraining factors influencing adoption. Interview data were analysed thematically. Conclusions, Expected Outcomes or Findings Preliminary analysis suggests that educators who report higher levels of confidence and familiarity with AI tend to engage with AI-supported formative assessment in more deliberate and pedagogically purposeful ways, particularly in relation to feedback provision, task design, and supporting student learning, rather than adopting such tools for lesson planning. However, emerging findings indicate that AI literacy alone does not fully account for patterns of adoption. Instead, a range of contextual factors including institutional assessment policies, workload and time pressures, access to professional development, ethical considerations, and peer support appear to play a significant role in shaping educators’ decisions to adopt or cautiously adapt AI-driven formative assessment tools. References Asian Development Bank. (2011, November). Higher education across Asia: An overview of issues and strategies (Report). Asian Development Bank. Bennett, R. E. (2011). Formative assessment: a critical review. Assessment in Education: Principles, Policy & Practice, 18(1), 5–25. https://doi.org/10.1080/0969594X.2010.513678 Black, P., & Wiliam, D. (1998). Assessment and classroom learning. Assessment in Education: Principles, Policy & Practice, 5(1), 7–74. https://doi.org/10.1080/0969595980050102 Chatfield, T. (2025). AI and the future of pedagogy (White paper). Sage. https://doi.org/10.4135/wp520172 Engeström, Y. (2001). Expansive Learning at Work: Toward an activity theoretical reconceptualization. Journal of Education and Work, 14(1), 133–156. https://doi.org/10.1080/13639080020028747 Hopfenbeck, T. N., Zhang, Z., Sun, S. Z., Robertson, P., & McGrane, J. A. (2023). Challenges and opportunities for classroom-based formative assessment and AI: A perspective article. Frontiers in Education, 8, 1270700. https://doi.org/10.3389/feduc.2023.1270700 Huesca, G., Elizondo-García, M. E., Aguayo-González, R., Aguayo-Hernández, C. H., González-Buenrostro, T., & Verdugo-Jasso, Y. A. (2025). Evaluating the Potential of Generative Artificial Intelligence to Innovate Feedback Processes. Education Sciences, 15(4), 505. https://doi.org/10.3390/educsci15040505 Simper, N., Mårtensson, K., Berry, A., & Maynard, N. (2022). Assessment cultures in higher education: reducing barriers and enabling change. Assessment & Evaluation in Higher Education, 47(7), 1016–1029. https://doi.org/10.1080/02602938.2021.1983770 Sulaiman, T., Kotamjani, S. S., Abdul Rahim, S. S., & Hakim, M. N. (2020). Malaysian public university lecturers’ perceptions and practices of formative and alternative assessments. International Journal of Learning, Teaching and Educational Research, 19(5), 379–394. https://doi.org/10.26803/ijlter.19.5.23 Wang, Y., Liu, C., & Tu, Y.-F. (2021). Factors Affecting the Adoption of AI-Based Applications in Higher Education: An Analysis of Teachers Perspectives Using Structural Equation Modeling. Educational Technology & Society, 24(3), 116–129. https://www.jstor.org/stable/27032860 09. Assessment, Evaluation, Testing and Measurement
Paper Dual-Track Governance in Teaching Quality Evaluation: Administrative Inspection and Peer Observation in China 1: University of Groningen, the Netherlands; 2: China University of Mining and Technology Presenting Author:Teaching quality evaluation has become a central instrument of contemporary education governance, yet many systems face an inherent tension. Accountability-oriented inspection and development-oriented professional learning serve fundamentally different purposes that are difficult to reconcile within unified evaluative frameworks (OECD, 2013). While some systems privilege one orientation over the other (e.g., the UK’s Ofsted inspections versus Finland’s trust-based professional autonomy), and others attempt integration within single instruments (e.g., the Danielson Framework in the United States), China adopts a distinctive alternative that has received limited systematic attention in international scholarship. This study examines the Chinese education system as a case in which administrative inspection (行政督导) and peer observation (同伴互助/听评课) are formally institutionalised as parallel, policy-mandated mechanisms operating within a unified national framework. Rather than treating these mechanisms as isolated practices, the study conceptualises teaching quality evaluation as a dual-track policy system shaped by multiple institutional pillars and enacted across governance levels. This configuration is particularly relevant given China’s vast teaching workforce and substantial regional variation in resources and governance capacity. Drawing on institutional theory, policy enactment theory, and cultural-historical perspectives, the study addresses four research questions: RQ1 (: How are administrative inspection and peer observation institutionally differentiated within China’s teaching quality evaluation policy system? (Differentiation) RQ2: Through which mechanisms are the two evaluative tracks coordinated while maintaining functional separation? (Coordination) RQ3: How is the dual-track framework recontextualised across governance levels and provincial contexts? (vertical and horizontal variation) Scott’s (2014) framework of institutional pillars provides the primary analytical lens for the current study. The regulative pillar operates through formal rules and sanctions that constrain behaviour and enforce compliance. The normative pillar functions through shared professional values and expectations about appropriate conduct. This framework maps productively onto China’s dual evaluation mechanisms: administrative inspection operates primarily through regulative channels (hierarchical authority, standardised procedures, rectification requirements), while peer observation draws on normative channels (collegial responsibility, collective improvement, mutual support). The concept of coupling (Meyer & Rowan, 1977; Coburn, 2004) offers analytical leverage for examining how the two tracks relate: whether loosely coupled, tightly integrated, or selectively connected through specific mechanisms. Policy enactment theory (Ball et al., 2012) complements our analysis by illuminating how policies are recontextualised as they travel across institutional levels. National frameworks establish structural templates, but provincial and school-level actors interpret, translate, and re-specify these designs according to local conditions (Braun et al., 2011). This perspective directly informs examination of cross-provincial variation in policy elaboration. The study contributes to international debates on teaching/teacher evaluation by advancing the concept of “differentiated parallel governance”, a policy configuration that institutionalises accountability and development as separate but coordinated tracks rather than attempting synthesis within unified instruments. This offers a distinctive response to the formative-summative tension that pervades teacher evaluation systems internationally. The current study provides a comparative reference point that expands the repertoire of available governance designs for European educators and policymakers navigating similar tensions between accountability demands and professional learning imperatives. Methodology, Methods, Research Instruments or Sources Used This study employs qualitative document analyses to examine the formal policy system of China’s dual-track teaching quality evaluation. The focus on basic education (primary and secondary schooling) reflects this sector’s governance under unified national policy frameworks. Policy documents serve as the core analytical unit, treating them as authoritative texts embodying specific institutional pillars and official discourses (Bowen, 2009). Documents were retrieved from official government databases including the National Laws and Regulations Database, Ministry of Education archives, and provincial education department (i.e., school) websites. The temporal scope (2010–2025) captures the contemporary evaluation framework established by China’s National Medium- and Long-Term Education Reform Plan (2010–2020) through recent developments. Initial retrieval yielded 167 potentially relevant documents. Application of inclusion criteria (official policy texts directly addressing teaching quality evaluation, specifying evaluation procedures or consequences) resulted in a final corpus of 96 documents distributed across three governance levels: national (n=14), provincial (n=61), and school (n=21). Seven provinces/municipalities were purposively selected to capture systematic variation in economic development, administrative capacity, and regional context: Beijing, Shanghai, Jiangsu, and Guangdong (economically developed with strong administrative capacity) and Guizhou, Gansu, and Xinjiang (less-developed contexts with more constrained resources). This contrast illuminates how identical national frameworks are differentially elaborated under varying economic and resource conditions. Analysis proceeded through four phases. First, corpus structuring systematically tagged documents across administrative hierarchy, evaluation track, regional affiliation, and chronological sequence. Second, thematic coding followed Saldaña's (2021) two-cycle approach, employing both deductive categories derived from the theoretical framework and inductive codes capturing emergent patterns. Inter-coder reliability, assessed through independent coding of 20 documents (21% of corpus), reached Cohen's kappa of 0.76. Third, cross-level alignment analysis compared national mandates with provincial and school-level texts to trace policy recontextualisation patterns. Fourth, cross-provincial comparison used analytical matrices to isolate regional variations within the unified national framework. This multi-level design enables systematic tracing of how the dual-track system is designed at the national level and recontextualised as it moves through the governance system, while cross-provincial comparison illuminates how differentiated elaboration emerges within shared structural frameworks. Conclusions, Expected Outcomes or Findings Institutional differentiation (RQ1): The findings reveal deliberate separation grounded in distinct institutional pillars. Administrative inspection operates through the regulative pillar: a vertically extended accountability chain where designated supervisors (positioned in policy texts as the "third eye" of education regulation) are authorised to inspect, demand rectification, and trigger formal consequences for schools and leaders. Peer observation, by contrast, is anchored in the normative pillar: horizontally organised teaching research groups (教研组) where colleagues engage in collaborative lesson study and reflective dialogue. Evaluative authority is distributed among peers rather than vested in external supervisors, with consequences flowing through collegial relationships rather than formal sanctions. Coordination mechanisms (RQ2): The two tracks are connected through selective and asymmetrical coupling rather than structural integration. Three mechanisms emerge: sequential linkage (inspection findings trigger targeted peer-based improvement activities), artefact circulation (shared teaching artefacts flow across governance routines without collapsing functional distinctions), and temporal complementarity (episodic inspection and continuous peer observation create layered evaluation). Explicit boundary-setting language in policy texts protects the developmental track from consequence-bearing use, preserving its normative foundation. Recontextualisation and variation (RQ3): Cross-provincial analysis reveals strong structural convergence alongside systematic variation in elaboration. All provinces uniformly adopt nationally mandated components, demonstrating coercive isomorphism. However, developed provinces (Beijing, Shanghai, Jiangsu, Guangdong) exhibit denser professional learning infrastructure, while less-resourced provinces (Guizhou, Gansu, Xinjiang) focus on procedural compliance and minimum standards. This suggests that genuine functional differentiation is achieved primarily in well-resourced contexts, carrying implications for equity in teacher professional learning opportunities. The study advances the concept of “differentiated parallel governance”. This offers an alternative to both integrated frameworks that risk accountability contamination of professional learning and siloed systems where tracks operate in isolation, with potential implications for European systems navigating similar tensions. References Ball, S. J., Maguire, M., & Braun, A. (2012). How schools do policy: Policy enactments in secondary schools. Routledge. Bowen, G. A. (2009). Document analysis as a qualitative research method. Qualitative Research Journal, 9(2), 27–40. https://doi.org/10.3316/QRJ0902027 Braun, A., Ball, S. J., Maguire, M., & Hoskins, K. (2011). Taking context seriously: Towards explaining policy enactments in the secondary school. Discourse: Studies in the Cultural Politics of Education, 32(4), 585–596. https://doi.org/10.1080/01596306.2011.601555 Coburn, C. E. (2004). Beyond decoupling: Rethinking the relationship between the institutional environment and the classroom. Sociology of Education, 77(3), 211–244. https://doi.org/10.1177/003804070407700302 Ehren, M. C. M., Altrichter, H., McNamara, G., & O'Hara, J. (2013). Impact of school inspections on improvement of schools. Educational Assessment, Evaluation and Accountability, 25(1), 3–43. https://doi.org/10.1007/s11092-012-9156-4 Meyer, J. W., & Rowan, B. (1977). Institutionalized organizations: Formal structure as myth and ceremony. American Journal of Sociology, 83(2), 340–363. OECD. (2013). Teachers for the 21st century: Using evaluation to improve teaching. OECD Publishing. https://doi.org/10.1787/9789264193864-en Saldaña, J. (2021). The coding manual for qualitative researchers (4th ed.). Sage. Scott, W. R. (2014). Institutions and organizations: Ideas, interests, and identities (4th ed.). Thousand Oaks, CA: Sage Publishing. 09. Assessment, Evaluation, Testing and Measurement
Paper Authentic Assessment and Academic Honesty in the Age of Artificial Intelligence: Assessing Spoken Argumentation through Socratic Seminars Nazarbayev Intellectual school of Science and Mathematics in Karagandy, Kazakhstan Presenting Author:Recent European policy documents emphasise competence-based education, formative assessment, transparency, and learner autonomy as key principles of assessment reform. At the same time, growing access to AI tools capable of generating written texts has intensified concerns about academic integrity and validity. These challenges extend beyond the European Union. In Kazakhstan, ongoing curriculum and digital reforms have led teachers to report similar difficulties in verifying authorship and assessing authentic understanding. This study addresses both local and European concerns and contributes to international discussions on assessment systems responding to AI. Recent research highlights the widespread use of AI among secondary and tertiary students and its implications for assessment design. Many students regularly use tools such as ChatGPT for writing, summarising, and idea generation, often without clear ethical guidance (Baig, 2025; HEPI, 2025). Studies in secondary education report increasing AI-assisted writing and problem-solving, alongside teacher uncertainty about detection and regulation (Eke, 2023). International organisations warn that detection technologies alone cannot ensure assessment validity and recommend redesigning tasks to prioritise reasoning, learning processes, and interaction (UNESCO, 2023; OECD, 2024). One response is renewed attention to spoken argumentation as both a learning goal and assessment format. Argumentation is recognised as central to critical thinking, disciplinary understanding, and civic participation (Kuhn, 1991; Osborne, 2010). Research shows that structured oral argumentation improves conceptual development and deeper understanding when students justify claims and respond to counterarguments (Nussbaum, 2023). Unlike writing, spoken argumentation occurs in real time, making reasoning visible and limiting reliance on artificial intelligence. Assessment theory supports the pedagogical value of dialogic and formative approaches. Black and Wiliam (2009) argue that assessment in classroom interaction enhances learning by making thinking explicit and enable timely feedback. Nicol and Macfarlane-Dick (2006) emphasise the importance of transparent criteria and peer assessment. From this perspective, assessment is viewed not as a tool for measurement, but as an integral part of instruction that shapes students’ understanding of quality and supports ethical academic practices. The Socratic seminar represents an appropriate format for integrating spoken argumentation into assessment. Grounded in dialogic traditions, Socratic seminars engage students in structured discussions around shared texts or concepts, encouraging interpretation, critique, and collaborative knowledge construction (Adler, 1982). Empirical studies report positive effects on students’ critical reading, reasoning, and oral communication, especially when discussions are supported through explicit scaffolding and reflective feedback (Billings & Fitzgerald, 2002). More recent research suggests that combining inner-circle participation with outer-circle observation and peer assessment allows teachers to observe both performance and process and increases students’ awareness of argumentative quality (Copeland, 2023). Argumentation skills can be applied across disciplines when supported by consistent instructional frameworks. In science education, structured argumentation supports scientific reasoning and conceptual understanding (Driver et al., 2000), while in language education oral discourse fosters interpretative and evaluative competence (Swain, 2006). However, relatively little research has examined the systematic development and assessment of spoken argumentation across subjects within a shared pedagogical design, particularly in relation to the emerging challenge of AI-related academic dishonesty. The present study seeks to address this gap by investigating an adapted Socratic seminar model implemented as an assessment tool in Grade 10 English and Biology classrooms in Kazakhstan. Research Questions The study is guided by the following two research questions:
Methodology, Methods, Research Instruments or Sources Used The study employed an exploratory action research design to investigate the development and assessment of spoken argumentation skills across subjects and to improve classroom practice through systematic inquiry. Action research was considered appropriate because it allowed the researcher-teachers to identify a pedagogical problem arising from everyday practice, implement a targeted intervention, and reflect on its effects within an authentic instructional context (Cresswell, 2013). Participants were eight Grade 10 learner groups comprising 96 students selected through convenience sampling as they were taught by the researcher-teachers during the same academic year at the Nazarbayev Intellectual School of Karaganda City. The shared institutional context implied comparable curricular requirements in English and Biology and enabled cross-group and cross-subject comparison while maintaining ecological validity. The intervention was implemented over one academic term following a cycle of problem exploration, planning, action, observation, and reflection. At the initial stage, difficulties were identified in both the development and reliable assessment of spoken argumentation, particularly in relation to clarity of claims, use of evidence, logical coherence, and responsiveness to counterarguments. Hence, an instructional strategy was adapted from the classical Socratic seminar model. The adaptation focused on scaffolding argument construction and engaging students in the co-creation of explicit assessment criteria for peer observation rubrics. During the action phase, the adapted Socratic seminar approach was implemented in all eight groups. In English lessons, students discussed literary texts, while in Biology lessons they analysed target scientific concepts. In each seminar, students in the inner circle participated in guided discussion, sharing interpretations and responding to peers, while students in the outer circle observed and conducted peer assessment using the agreed rubrics. A mixed-methods approach to data collection was adopted. Quantitative data were obtained from rubric-based scores on pre- and post-intervention speaking tasks. Descriptive statistics were used to examine changes in mean scores across groups and subjects. The overall mean score increased from 11.4 to 15.8, showing an average improvement of 4.4 points. The biggest increase was observed in the use of evidence (from 2.6 to 4.3) and reasoning (from 2.4 to 4.0). Qualitative data were collected through teacher observation notes and focus-group discussions. Qualitative analysis revealed that 18 out of 24 students reported increased confidence in speaking, and 15 students indicated that the rubric made assessment criteria clearer. The qualitative data were coded thematically into three main themes: development of argument structure, increased learner awareness of assessment criteria, and implementation challenges. Conclusions, Expected Outcomes or Findings The present study demonstrates the potential of an adapted Socratic seminar approach to support both the development and assessment of spoken argumentation skills across English and Biology and responding to the emerging challenge of AI-related academic dishonesty. The findings indicate that transparent criteria, peer assessment, and structured dialogic interaction can enhance students’ ability to formulate clear claims, use supporting evidence, and engage more confidently in academic discussion. Importantly, similar patterns of improvement were observed across subjects, suggesting that argumentation skills are transferable when consistent instructional and assessment frameworks are applied. From a methodological perspective, the study confirms the value of action research as a means of improving assessment practice in authentic classroom contexts. The combination of quantitative and qualitative data provided a detailed picture of both learning outcomes and classroom processes and highlighted the importance of triangulation when evaluating complex oral competencies. At the same time, the challenges reported by teachers, particularly the difficulty of facilitating discussion while conducting reliable assessment in time-constrained settings, underscore the need for further refinement of assessment tools, increased use of peer and self-assessment, and professional development focused on dialogic evaluation strategies. At a broader level, the study contributes to European and international debates on assessment reform in the age of artificial intelligence. Rather than relying on detection technologies or restrictive policies, the findings support a pedagogical shift towards assessment designs based on reasoning, interaction, and process. Spoken argumentation offers a sustainable response that not only reduces opportunities for AI-generated work but also strengthens core competences, such as critical thinking, communication, and learner autonomy. References Adler, M. J. (1982). The Paideia proposal: An educational manifesto. New York: Macmillan. Baig, M. I., & Yadegaridehkordi, E. (2025). Factors influencing academic staff satisfaction and continuous usage of generative artificial intelligence (GenAI) in higher education. International Journal of Educational Technology in Higher Education, 22(1), 5. Billings, L., & Fitzgerald, J. (2002). Dialogic discussion and the Paideia seminar. American Educational Research Journal, 39(4), 907–941. Black, P., & Wiliam, D. (2009). Developing the theory of formative assessment. Educational Assessment, Evaluation and Accountability, 21(1), 5–31. Nussbaum, E. M., Dove, I. J., & Putney, L. G. (2023). Bridging Dialogic Pedagogy and Argumentation Theory through Critical Questions. Dialogic Pedagogy, 11(3). Copeland, M. (2023). Socratic circles: Fostering critical and creative thinking in middle and high school. Routledge. Cresswell, J. (2013). Qualitative inquiry & research design: Choosing among five approaches. Driver, R., Newton, P., & Osborne, J. (2000). Establishing the norms of scientific argumentation in classrooms. Science Education, 84(3), 287–312. HEPI. (2025). Student generative AI survey 2025. Oxford: Higher Education Policy Institute. Kuhn, D. (1991). The skills of argument. Cambridge: Cambridge University Press. Eke, D. O. (2023). ChatGPT and the rise of generative AI: Threat to academic integrity?. Journal of Responsible Technology, 13, 100060. Nicol, D., & Macfarlane-Dick, D. (2006). Formative assessment and self‐regulated learning. Studies in Higher Education, 31(2), 199–218. OECD. (2024). Artificial intelligence in education: Challenges and opportunities. Paris: OECD. Osborne, J. (2010). Arguing to learn in science. Science, 328(5977), 463–466. Swain, M. (2006). Languaging, agency and collaboration in advanced second language proficiency (pp. 95-108). UNESCO. (2023). Guidance for generative AI in education and research. Paris: UNESCO. | ||