Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
|
Daily Overview |
| Session | ||
31 SES 03 B JS: Language, AI & Assessment - Joint session (NW 31 and Nw 09)
Joint Paper Session | ||
| Presentations | ||
31. LEd – Network on Language and Education
Paper Assessing Authentic Language Production in the Age of AI Writing Tools in IB English B HL. Nazarbayev Intellectual schools, Kazakhstan Presenting Author:Proposal Information (461) Recent developments in generative AI writing tools are radically changing how second language learners prepare, write and revise their texts. More specifically, in EFL, they are able to quickly refine grammar, structure and clear expression of meaning by producing instant personalised feedback and automatic suggestions (Dong, 2023). But this help introduces a new assessment challenge: in other words, if AI helps shape most of the final text, does that text still embody the learner's own language proficiency? This conflict is especially salient in high stakes international arenas like IB English B HL where the aim of assessment is to reflect independent communicative competence and real voice. Studies with EFL students suggest that, though AI-assisted assessment can counter linguistic uncertainty and enhance overall quality, over-reliance may result in over-harmonisation of texts and a loss of culturally situated individual expression (Werdiningsih, Marzuki & Rusdin, 2024). Even students have stressed the importance of mixing AI guidance with their own judgment to keep the originality and authenticity (Werdiningsih et al., 2024). From a language assessment standpoint, generative AI undermines fundamental notions of validity, fairness and reliability. In a recent systematic review, it has been reported the risks of bias, inequality of access, and unequal impact along large-scale, local and classroom levels of assessment if AI is introduced without scrutiny (Chuang & Yan, 2025). Simultaneously academic integrity frameworks currently do not effectively classify and regulate the multitude of AI-based writing tools currently accessible in education, leading to a potential grey area as opposed to evident instances of violation (Roe, Renandya & Jacobs, 2023). Instead of treating AI as a threat to humans alone, a range of researchers advocate for a participatory, human-AI model, in which teachers are at the heart as interpreters, guides and ethical gatekeepers of evaluation (Amin, 2023). This raises the question of what constitutes “authentic” writing within digitally mediated settings. This study therefore asks: In what way does the deployment of AI software tools challenge the authenticity of language writing assessment within IB English B HL, and how assessment could change in order to maintain relevance and legitimacy on international stage? The study is grounded on sociocultural theory of mediated learning, which sees AI as a medium that enhances but also redistributes cognitive efforts. For this reason, authenticity is defined as a state of learner choice in how and when AI is offered input rather than one that reflects a lack of tools. For that, in IB classes — a place in which students have different linguistic resources — preserving voice becomes even more crucial (Werdiningsih et al., 2024). Through examining AI mediation in novel ways beyond the framework of AI vs. no-AI writing, the project endeavours to provide an applicable evidence for assessment policy and practice in generative AI era (Chuang & Yan, 2025). Methodology, Methods, Research Instruments or Sources Used Methodology (193) A mixed-methods classroom study consisted of two study groups. Under 3 conditions, students completed similar writing tasks: No AI support. AI employed for revision and for feedback purposes only. Unrestricted AI assistance. Texts, drafts and histories of revisions were collected in order to trace the use of AI suggestions. The comparative linguistic analyses showed the gains in accuracy and complexity that AI feedback (Dong, 2023) typically drives, and the negative differences of individual-sounding voice that can deteriorate under heavy AI mediation (Werdiningsih et al., 2024). Students were required to write reflective comments on their decisions regarding accepting and rejecting AI suggestions, thus encouraging a critical AI literacy developed in recent EFL research (Werdiningsih et al., 2024; Amin, 2023). Semi-structured interviews explored perceptions of authenticity and ownership. To assess assessment validity and fairness, teachers scored scripts against IB criteria and then evaluated anonymised or minimally AI-edited versions to find out how slick AI-mediated form affects judgment (Chuang & Yan, 2025). Study groups comparisons may alleviate concerns of inequity in assessment that may arise from not having access to or familiarity with AI tools (Chuang & Yan, 2025; Roe et al., 2023). Conclusions, Expected Outcomes or Findings Conclusion (225) Anticipated findings or results. It is anticipated that the use of AI will lead to considerably enhanced surface accuracy and coherence that support well-reported pedagogical advantages of AI-mediated writing (Dong, 2023; Amin, 2023). Nevertheless, heavy mediation of AI is also likely to be associated with lower levels of indicators of personal voice and culturally inflected expression, consistent with issues of authentic student reports (Werdiningsih et al., 2024). Teacher scores are anticipated to be higher in AI-assisted texts, though part of this gain might depend on refined form as opposed to more profound communicative competence, making it likely that product-only assessment raises questions surrounding validity (Chuang & Yan, 2025). Findings may also disclose inequities associated with different AI access and skill, strengthening demands for more transparent, educative integrity-practice policies as opposed to purely punitive regulations (Roe et al., 2023). The research posits an argument for a shift away from exclusion of AI to the assessment of human–AI interaction as responsible and agentive. Practical implications are incorporating process evidence (drafts and reflections), teaching critical self-reflection on AI feedback in a public way (Amin, 2023; Werdiningsih et al., 2024) and creating clear international standards for ethical use of AI in language assessment (Chuang & Yan, 2025). In this manner, authenticity is redefined as educated, autonomous meaning making in conversation with digital technologies, rather than their total absence. References References: Amin, M. Y. M. (2023). AI and ChatGPT in language teaching: Enhancing EFL classroom support and transforming assessment techniques. Chuang, P.-L., & Yan, X. (2025). Language assessment in the era of generative artificial intelligence: Opportunities, challenges, and future directions. Dong, Y. (2023). Revolutionizing academic English writing through AI-powered pedagogy: Practical exploration of teaching process and assessment. Roe, J., Renandya, W. A., & Jacobs, G. M. (2023). A review of AI-powered writing tools and their implications for academic integrity in the language classroom. Werdiningsih, I., Marzuki, & Rusdin, D. (2024). Balancing AI and authenticity: EFL students’ experiences with ChatGPT in academic writing. Carvalho L., Martinez-Maldonado R., Tsai Y.-S., Markauskaite L., and De Laat M., (2022) How can we design for learning in an AI world?, Computers and Education: Artificial Intelligence. He, J. (2020). The impact of AI-powered apps on young learners' language acquisition. Educational Technology Research and Development, 68(3), 1475-1487. Zhang, L., Zhao, D., & Yang, X. (2018). Using chatbots in education: Application scenarios, impacts, and trends. Educational Technology Research and Development, 66(6), 1463- 1488. 31. LEd – Network on Language and Education
Paper Validating AI-Enhanced Criteria-Based Assessment of Pre-Service Language Teachers' Intercultural Communicative Competence: A Human-AI Collaboration Framework for Multilingual Contexts Kazakh Ablai Khan University of International Relations and World Lang, Kazakhstan Presenting Author:This study addresses the following main research question: to what extent do assessment tools improved by artificial intelligence demonstrate validity and reliability in assessing the intercultural communicative competence of foreign language teachers (ICC) prior to employment, compared with human expert evaluators, and how can human-artificial intelligence collaboration optimize the assessment process in multilingual contexts? Additional research questions are the following: 1. What is the degree of mutual agreement between the assessments of ICC obtained using artificial intelligence and the expert assessments of teachers before starting training using criteria-based headings? 2. Which model of human-artificial intelligence interaction ensures maximum assessment reliability while maintaining the professional freedom of action of the teacher in a multilingual educational context? This research integrates four complementary theoretical frameworks to conceptualize and validate AI-enhanced assessment of ICC in pre-service teacher education. Methodology, Methods, Research Instruments or Sources Used This study uses a mixed approach to collect and analyze quantitative and qualitative data simultaneously to verify the results of the ICC assessment using AI. This approach allows for real-time comparison of AI assessments, expert assessments, and teacher opinions, providing comprehensive evidence of validity. The study will involve 100 students studying to become a foreign language teacher from Yesenov University in Aktau, Kazakhstan. Targeted sampling provides a variety of linguistic and cultural experiences. Participants are enrolled in the sixth semester (3rd year of study) of the teacher training program with completed coursework on intercultural communication. Data Collection Instruments Quantitative Strand: ICC Performance Assignments Three real-world intercultural communication scenarios—video-recorded role-plays, written reflective analyses, and asynchronous cross-cultural discussion facilitation—are completed by participants. Byram's five savoirs serve as the basis for task creation, and criteria-based rubrics (1–5 scale spanning 12 descriptors) are used for scoring. AI Assessment Tool: A custom-developed Natural Language Processing (NLP) system merging automated essay scoring for written reflections and multimodal analysis for video performances. Scores produced by the program are in line with the same 12-descriptor criteria that human raters use. Perceived Usefulness, Perceived Ease of Use, Attitude, Digital Literacy, and Behavioral Intention are all measured by this modified TAM test (32 items, 5-point Likert scale). Qualitative Strand: Semi-structured Interviews (n = 30): 45-minute interviews examining opinions regarding the usefulness, fairness, transparency, and influence on professional identity of AI assessments. Think-Aloud Protocols (n=20): Participants discuss how accurate and helpful they believe their AI-generated comments to be. Analysis of Data Quantitative Evaluation: Inter-rater reliability: Weighted Cohen's Kappa coefficients comparing AI-generated scores to consensus scores from three experienced human raters (κ > 0.70 indicates acceptable agreement). Pearson correlations between AI and human aggregate scores are used in correlation analysis. TAM validation: Relationships between Perceived Usefulness, Perceived Ease of Use, Attitude, and Behavioral Intention are tested using structural equation modeling (SEM). Qualitative Analysis: Thematic analysis: NVivo software is used to inductively code interview transcripts and think-aloud protocols in order to find themes pertaining to professional agency, trust, transparency, and justice. Cohen's Kappa is used to evaluate inter-coder reliability (two coders). Integration: Data convergence using side-by-side comparison matrices that show both qualitative themes about perceived validity and quantitative reliability statistics. Ethical Considerations The focus of informed consent is on GDPR compliance, voluntary involvement, and data confidentiality. Participants are still able to examine assessments produced by AI and ask for a human reevaluation. Every piece of information will be saved on encrypted servers using pseudonyms. Conclusions, Expected Outcomes or Findings Two areas of educational research are expected to advance as a result of this study. In the first place, it will offer empirical validation of AI-enhanced assessment for culturally situated competencies, expanding the scope of existing AI assessment research, which mostly concentrates on linguistic accuracy, to the more intricate field of intercultural communicative competence. Second, the study is anticipated to produce a theoretically grounded human-AI collaboration model for teacher education assessment, outlining the best way to divide tasks between human judgment (interpretation of cultural appropriateness, critical awareness) and automated scoring (pattern recognition in linguistic and behavioral indicators). Anticipated results include: 1. Byram's ICC model for AI-enhanced scoring has been operationalized through validated assessment rubrics that are applicable to all teacher education programs. 2. Evidence-based recommendations for the distribution of tasks between humans and AI, outlining which ICC dimensions are better served by automation as opposed to professional judgment. 3. Teacher acceptability is predicted by TAM-validated factors (perceived usefulness > perceived ease of use, moderated by digital literacy). 4. GDPR compliance, explainability standards, and bias prevention for multilingual populations are among the implementation recommendations. References 1 Byram, M. (1997). Teaching and assessing intercultural communicative competence. Multilingual Matters. 2 Creswell, J. W., & Plano Clark, V. L. (2017). Designing and conducting mixed methods research (3rd ed.). SAGE Publications. 3 Davis, F. D. (1989). Perceived usefulness, perceived ease of use, and user acceptance of information technology. MIS Quarterly, 13(3), 319–340. https://doi.org/10.2307/249008 4 Deardorff, D. K. (2006). Identification and assessment of intercultural competence as a student outcome of internationalization. Journal of Studies in International Education, 10(3), 241–266. https://doi.org/10.1177/1028315306287002 5 European Commission. (2022). Digital Competence Framework for Educators (DigCompEdu). Publications Office of the European Union. https://doi.org/10.2760/159770 6 Gabon, D. C., et al. (2025). Automated grading of essay using natural language processing: A comparative analysis with human raters across multiple essay types. Journal of Information Systems Engineering and Management, 10(6), Article 700. https://doi.org/10.55267/iadt.07.14700 7 Grand-Guillaume-Perrenoud, J. A., et al. (2023). Mixed methods instrument validation: Evaluation procedures integrating the quality criteria of congruence, convergence, and credibility. BMC Medical Research Methodology, 23(1), Article 22. https://doi.org/10.1186/s12874-022-01811-1 8 Kong, S. C., et al. (2024). Examining teachers' behavioural intention of using generative AI: A modified technology acceptance model. Computers and Education: Artificial Intelligence, 6, Article 100210. https://doi.org/10.1016/j.caeai.2024.100210 9 Ministry of Digital Development, Innovation and Aerospace Industry of the Republic of Kazakhstan. (2024). Concept for the Development of Artificial Intelligence in the Republic of Kazakhstan for 2024-2029. Government Publishing. 10 Pack, A., et al. (2024). Large language models and automated essay scoring of second language writing. Computers and Education: Artificial Intelligence, 6, Article 100353. https://doi.org/10.1016/j.caeai.2024.100353 11. Tariq, M., et al. (2024). Evaluating human-AI collaboration: A review and methodological framework. arXiv preprint arXiv:2407.19098. https://doi.org/10.48550/arXiv.2407.19098 12 UNESCO. (2024). AI competency framework for teachers. UNESCO Publishing. https://unesdoc.unesco.org/ark:/48223/pf0000390389 31. LEd – Network on Language and Education
Paper Reconceptualizing Assessment Competence for Secondary EFL Teachers in the Generative AI Era: A Continuum-Based Framework Beijing Normal University;Chongqing Yuzhong Institute for Teacher Education, PRC Presenting Author:The digital transformation of education, accelerated by the integration of generative AI (GenAI), has fundamentally reshaped language assessment practices. Secondary EFL teachers now face the dual challenge of navigating digitally mediated evaluation while addressing discipline-specific pedagogical, ethical, and technological demands (Cui et al., 2025; UNESCO, 2024). Existing competence frameworks often adopt a fragmented view, focusing either on generic digital literacy or discrete assessment skills, thereby overlooking the integrated and contextual nature of teacher competence in AI-enhanced environments (Estaji et al., 2024). Grounded in the “competence as a continuum” perspective (Blömeke et al., 2015), this study conceptualizes teacher competence as a multidimensional construct spanning dispositions, situation-specific skills, and observable performance. In the GenAI era, this continuum must also encompass emerging competencies—such as prompt design, critical evaluation of AI-generated outputs, and human–AI collaborative feedback—which are not yet systematically represented in language assessment literacy models (Kremmel & Harding, 2020; TALiGAI framework). This study therefore addresses the following research question: How can a continuum-based digital assessment competence framework for secondary EFL teachers be developed and validated to incorporate both established dimensions and AI-responsive capabilities? This study aims to develop and validate an integrated digital assessment competence framework for secondary EFL teachers that aligns with a continuum-based understanding of teacher competence and responds to the evolving demands of GenAI. The framework synthesizes established dimensions of language assessment literacy—such as those identified in empirical models like the Language Assessment Literacy Survey (Kremmel & Harding, 2020)—with forward-looking, AI-integrated competencies highlighted in recent scoping reviews (e.g., knowledge of AI ethics, critical appraisal of AI outputs, adaptive pedagogical practices). It intends to provide a structured, discipline-sensitive tool to guide teacher professional development, curriculum design, and reflective assessment practice in digitally enhanced language learning environments. Methodology, Methods, Research Instruments or Sources Used The study employs a sequential mixed-methods design comprising three phases, informed by instrument-development approaches in language assessment literacy research (Kremmel & Harding, 2020): Systematic analysis of existing assessment literacy frameworks(e.g., DC4LT, TALiGAI) and digital competence frameworks (e.g., DigCompEdu, China’s Teacher Digital Literacy standards) and empirical literature on GenAI in education to derive an initial set of competency dimensions. In-depth behavioral event interviews with 30 secondary EFL teachers to contextualize and enrich the framework based on real-world instructional experiences, challenges, and adaptive strategies in AI-enhanced assessment. An iterative Delphi process involving 15 experts in language assessment, digital pedagogy, and EFL teacher education to refine, consolidate, and validate the framework for disciplinary relevance, practical applicability, and alignment with continuum-based competence theory. Conclusions, Expected Outcomes or Findings The study will yield a validated, multidimensional competence framework organized along five interrelated dimensions—Digital Assessment Cognition, Attitude, Skill, Behavior, and Ethics—each conceptualized as part of a dynamic continuum from dispositional traits to situated performance. The framework will explicitly integrate AI-related competencies as cross-cutting components across these dimensions, reflecting the need for teachers to engage critically and creatively with GenAI in assessment contexts (e.g., designing authentic tasks, ensuring academic integrity, fostering student AI literacy). By synthesizing continuum theory (Blömeke et al., 2015) with empirically grounded dimensions of language assessment literacy and insights from the TALiGAI framework, this study offers a robust foundation for EFL teacher education, professional development programs, and the promotion of ethically informed, technologically adaptive assessment practices in the age of AI. References Blömeke, S., Gustafsson, J.-E., & Shavelson, R. J. (2015). Beyond dichotomies: Competence viewed as a continuum. Zeitschrift für Psychologie, 223 (1), 3–13. Cui Y, Meng Y, Tang L. Reconsidering teacher assessment literacy in GenAI-enhanced environments: A scoping review[J]. Teaching and Teacher Education, 2025, 165: 105163. Estaji, M., Banitalebi, Z., & Brown, G. T. (2024). The key competencies and components of teacher assessment literacy in digital environments: A scoping review. Teaching and Teacher Education, 141, 104497. Kremmel, B., & Harding, L. (2020). Towards a comprehensive, empirical model of language assessment literacy across stakeholder groups: Developing the Language Assessment Literacy Survey. Language Assessment Quarterly, 17(1), 100–120. UNESCO. (2024). AI Competency Framework for Teachers. UNESCO Publishing. | ||