Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
Please note that all times are shown in the time zone of the conference. The current conference time is: 19th Aug 2026, 20:15:34 EET
|
Daily Overview |
| Session | ||
12 SES 06 A JS: Joint Symposium - From Log Data to Research Data: Quality, Documentation, Reuse, and Protection in Large-Scale Educational Studies - NW 09, NW 12 and NW 27
Joint Symposium Session
| ||
| Presentations | ||
12. Open Research in Education
Symposium From Log Data to Research Data: Quality, Documentation, Reuse, and Protection in Large-Scale Educational Studies Digitalization has profoundly reshaped data production in educational research. Technology-based assessments now generate large volumes of log-files holding process-oriented data, while at the same time expectations regarding transparency, reproducibility, and data reuse have increased across the social sciences. Against this background, educational researchers face a dual challenge: managing technically complex data such as log files and embedding these data within broader infrastructures and governance frameworks that enable responsible sharing and reuse. This symposium brings together four complementary contributions that address these challenges across different stages of the research data lifecycle. Three of the contributions focus specifically on log data from technology-based large-scale assessments, examining issues of documentation, data quality, and data protection. The fourth contribution deliberately broadens the perspective by examining educational research data more generally within a national social science repository. Together, the presentations situate log data within the wider ecosystem of educational research data practices rather than treating them as a standalone methodological innovation. The first contribution, by Breitwieser addresses the challenge of making log data reusable through adequate documentation. Log data are inherently context-dependent: Individual events can only be interpreted in relation to the digital system that generated them. Documentation is crucial in helping secondary researchers to understand this relation. Drawing on examples from international large-scale assessments and digital learning environments, the presentation illustrates why conventional documentation formats are often insufficient. It compares alternative strategies, such as enhanced system descriptions and log-data replays, and concludes with practical recommendations and proposed minimum standards for documenting log data in a way that supports transparency, reproducibility, and secondary use. The second contribution, by Borisova, examines data protection challenges associated with sharing and reusing log data under the General Data Protection Regulation (GDPR). With the shift to online testing, large-scale assessments increasingly collect detailed behavioral traces that may constitute personal data even after direct identifiers are removed. Drawing on experiences from TIMSS 2023, the presentation discusses challenges related to transparency, purpose limitation, and re-identification risk. It illustrates how legal requirements can be addressed through improved participant information, technical safeguards, and contractual measures, while still enabling scientifically meaningful data access. The third contribution, by Sibberns, critically examines the widespread assumption that machine-generated data are inherently clean and error-free. Based on empirical experiences from ILSA log-file processing, the presentation demonstrates how technical failures, programming inconsistencies, and administrative deviations introduce systematic errors into log data. It discusses methods for identifying such issues and compares strategies for handling them, emphasizing that preprocessing decisions have substantive analytical consequences and must therefore be explicitly documented. The fourth contribution, by Heers, shifts the focus from specific data types to research data infrastructures. Using metadata from SWISSUbase, Switzerland’s national social science data repository, the presentation profiles educational research datasets in terms of data types, institutional affiliations, and regional distribution. These patterns are compared with publicly available information on education research outputs to identify gaps and areas of underrepresentation. The analysis highlights structural and cultural barriers that continue to limit data sharing in education research and discusses how repositories can address these challenges through tailored support, guidance on ethical and practical issues, and incentives aligned with FAIR principles. Taken together, the symposium highlights that the scientific value of both log data and educational research data more broadly depends on coordinated attention to data quality, documentation, legal compliance, and infrastructural support. By combining log-data–specific perspectives with a repository-based analysis, the symposium aims to foster a more integrated and reflexive approach to research data practices in large-scale educational studies. References Becker, B., Neuendorf, C., & Jansen, M. (2022). Nutzung von Logdaten in der empirischen Bildungsforschung—Eine Bedarfsanalyse. https://doi.org/10.5281/ZENODO.7030996 Chen, G., Liu, Y., & Mao, Y. (2024). Understanding the log file data from educational and psychological computer-based testing: A scoping review protocol. PLoS ONE, 19(5), e0304109. European Union. (2016). General Data Protection Regulation (EU) 2016/679. Goldhammer, F., Hahnel, C., Kroehne, U., & Zehner, F. (2021). From byproduct to design factor: On validating the interpretation of process indicators based on log data. Large-Scale Assessments in Education, 9(1), 1–25. Heers, M. (2023). Data sharing in the social sciences. FORS Guides. Kroehne, U., & Goldhammer, F. (2025). Software tools for analyzing log data. In Innovative Digital-Based International Large-Scale Assessments. Springer. Late, E., & Ochsner, M. (2024). Re-use of research data in the social sciences. PLOS ONE, 19(5), e0303190. Wilkinson, M. D., et al. (2016). The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data, 3, 160018. Presentations of the Symposium Making Log Data Reusable: Documentation Challenges and Recommendations
Log data (i.e., event data documenting individuals’ interactions with a digital environment) have become an important source of information in empirical educational research (Becker et al., 2022; Goldhammer et al., 2021). They provide fine-grained insights into learners’ behavior and performance-related processes, thus extending beyond what can be captured with traditional questionnaire or test score data. However, log data also pose specific challenges for data sharing and reuse. Their high volume, technical nature, and strong context dependence make it difficult to implement the FAIR principles (Wilkinson et al., 2016), in particular with respect to reusability. Consequently, log data remain underutilized in secondary analyses despite their high scientific value and reuse potential (Becker et al., 2022).
One of the central challenges in making log data reusable is the provision of high-quality documentation. Log data typically cannot be interpreted without extensive background knowledge, much of which is held implicitly by the primary researchers. This concerns not only the structure of the log data itself but also the digital system that generated them. The meaning of a log event, such as a mouse click, can only be understood in relation to the technical elements of the system, including which interface elements can be interacted with and what is displayed on the screen at the time of the interaction. Codebooks and purely textual descriptions are therefore often insufficient to convey this information.
In addition, log data typically undergo a multi-step processing procedure to derive interpretable indicators of learners’ behavior and performance-related processes (Goldhammer et al., 2021). This procedure includes data preprocessing, the extraction of low-level features, and the modeling of process indicators (Kroehne & Goldhammer, 2025). Without transparent documentation of these steps, the interpretability, reproducibility, and scientific reuse of the data are severely limited. Taken together, these challenges highlight the need to develop high-quality log data documentation strategies.
Our presentation highlights these challenges and illustrates practical approaches to addressing them. Our examples draw on documentation strategies from large-scale assessments such as PIAAC (OECD, 2013), PISA (OECD, 2016), and TIMSS (Martin et al., 2020), as well as our own experience documenting log data from digital learning environments. For instance, we compare different methods for documenting the relationship between log data and the underlying digital environment, ranging from static documentation formats to log data replays, and discuss their respective strengths and limitations. The presentation concludes with concrete recommendations and proposed minimum requirements for documenting log data.
References:
Becker, B., Neuendorf, C., & Jansen, M. (2022). Nutzung von Logdaten in der empirischen Bildungsforschung—Eine Bedarfsanalyse. https://doi.org/10.5281/ZENODO.7030996
Goldhammer, F., Hahnel, C., Kroehne, U., & Zehner, F. (2021). From byproduct to design factor: On validating the interpretation of process indicators based on log data. Large-Scale Assessments in Education, 9(1), 20. https://doi.org/10.1186/s40536-021-00113-5
Kroehne, U., & Goldhammer, F. (2025). Software Tools for Analyzing Log Data. In L. Khorramdel, M. Von Davier, & K. Yamamoto (Eds.), Innovative Digital-Based International Large-Scale Assessments (pp. 657–691). Springer Nature Switzerland. https://doi.org/10.1007/978-3-031-90951-1_27
Martin, M. O., von Davier, M. & Mullis, I. V. (2020). Methods and Procedures: TIMSS 2019 Technical Report. TIMSS & PIRLS International Study Center.
OECD. (2013). Technical Report of the Survey of Adult Skills PIAAC (Second Edition). OECD Publishing.
OECD. (2016). PISA 2015 assessment and analytical framework. OECD Publishing. https://doi.org/10.1787/19963777
Wilkinson, M. D., Dumontier, M., Aalbersberg, Ij. J., Appleton, G., Axton, M., Baak, A., Blomberg, N., Boiten, J.-W., da Silva Santos, L. B., Bourne, P. E., Bouwman, J., Brookes, A. J., Clark, T., Crosas, M., Dillo, I., Dumon, O., Edmunds, S., Evelo, C. T., Finkers, R., … Mons, B. (2016). The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data, 3(1), Article 1. https://doi.org/10.1038/sdata.2016.18
Data Protection Challenges in the Sharing and Re-Use of Log Data in Large-Scale Educational Assessments: Insights from TIMSS 2023
With the shift to the online testing environment, the International Large-Scale Assessments (ILSAs) started to collect new category of personal data, namely log data (Case C-579/21), recording how participants interact with test platforms, which includes capturing clicks, keystrokes, time spent (Huber, K., Bannert, M.), navigation paths, and even pauses.
While log data offers many benefits from predicting outcomes and design interventions (Arizmendi, C.J., Bernacki, M.L., Raković, M. et al.) to supporting research and innovation, the sharing and reuse of captured log data raise important data protection considerations (Aghili, R., Li, H., & Khomh, F. (2024).
This contribution will focus on some of the data protection challenges under the General Data Protection Regulation (GDPR) when it comes to sharing and re-using log data in educational research and how IEA has addressed these challenges in TIMSS 2023.
The first challenge relates to the principle of transparency (Article 5(1)(a) GDPR), including the data subject’s right to be properly informed about the processing of their personal data (Articles 12-14 GDPR). This can be problematic, because for various reasons, which will be discussed, study participants might not be adequately informed about the types of log data collected, how long log data will be stored, or whether log data may be shared with third parties or reused for future research.
The second challenge relates to the principle of purpose limitation (Article 5(1)(b) GDPR), which states that personal data must be collected for specified and explicit purposes. Further processing of personal data is allowed for archiving purposes in the public interest, scientific or historical research or statistical purposes under Article 89(1), subject to appropriate safeguards and only if processed in a manner that is compatible with the initial purposes.
The third challenge relates to the ever-present re-identification risk. Even when direct identifiers are removed, log data can remain personal data under the GDPR, as detailed behavioral traces may allow individuals to be singled out, especially when combined with contextual information or external datasets.
The presentation concludes with insights from TIMSS 2023, showing how IEA strives to address those challenges via data protection declarations, whereby participants are informed about the processing of their personal data, as well as by implementing various safeguards to reduce the risk of re-identification, which are both technical and contractual in nature.
References:
Case C-579/21, J.M. v Pankki S (CJEU, 22 June 2023) EU:C:2023:501
Huber, K., Bannert, M. Investigating learning processes through analysis of navigation behavior using log files. J Comput High Educ 36, 683–701 (2024), p684. https://doi.org/10.1007/s12528-023-09372-3
Arizmendi, C.J., Bernacki, M.L., Raković, M. et al. Predicting student outcomes using digital logs of learning behaviors: Review, current standards, and suggestions for future work. Behav Res 55, 3026–3054 (2023), p3026. https://doi.org/10.3758/s13428-022-01939-9
Aghili, R., Li, H., & Khomh, F. (2024). Protecting Privacy in Software Logs: What Should Be Anonymized? ArXiv. https://arxiv.org/abs/2409.11313, p FSE060:2
European Union (2016). General Data Protection Regulation (EU) 2016/679.
The Illusion of Clean Data: Reflections on Log-File Processing in ILSA Studies
Large-scale assessments increasingly rely on computer-generated process data to complement traditional test responses. Log files from computer-based assessments are often treated as a particularly reliable data source, based on the implicit assumption that machine-generated data is inherently precise, complete, and free from error. This assumption, however, obscures a range of data quality issues that can substantially affect secondary analyses if left unexamined. In the context of International Large-Scale Assessments (ILSAs), where comparability and reproducibility are central concerns, such issues warrant closer scrutiny.
This presentation challenges the notion of “clean” log data by drawing on empirical experiences from processing log files in ILSA studies. It demonstrates that log data frequently contain errors that originate not from respondents, but from the technical and administrative infrastructure of computer-based testing. These errors may arise from programming flaws in assessment software, inconsistencies across software versions, hardware or operating system failures, unstable network connections, or deviations from prescribed testing procedures at the field level. While often rare in relative terms, such errors can have disproportionate consequences for process indicators, time-based measures, or sequence analyses.
Using selected real-world examples, the presentation illustrates typical error patterns found in ILSA log files. These include missing or duplicated events, implausible time stamps, corrupted event sequences, and inconsistencies between log data and background or response data. The presentation discusses how such issues can be detected through systematic plausibility checks, rule-based validation, and cross-referencing with technical documentation and field reports. Particular attention is paid to distinguishing genuine respondent behavior from artefacts introduced by the testing system itself.
Beyond identification, the presentation addresses strategies for handling erroneous log data. It compares different approaches, ranging from data exclusion and flagging to targeted correction and imputation, and discusses their respective implications for transparency, replicability, and substantive interpretation. The presentation argues that decisions made at the log-processing stage are analytically consequential and should therefore be explicitly documented rather than treated as purely technical preprocessing steps.
In conclusion, the presentation advocates for a more critical and reflexive approach to log-file processing in ILSA research. Recognizing and addressing the illusion of clean data is essential not only for improving data quality, but also for strengthening the methodological foundations of process data research in large-scale assessments.
References:
Chen G, Liu Y, Mao Y (2024) Understanding the log file data from educational and psychological computer-based testing: A scoping review protocol. PLoS ONE 19(5): e0304109. https://doi.org/10.1371/journal.pone.0304109
Goldhammer, F., Hahnel, C., Kroehne, U., & Zehner, F. (2021). From byproduct to design factor: On validating the interpretation of process indicators based on log data. Large-Scale Assessments in Education, 9(1), 1–25. https://doi.org/10.1186/s40536-021-00113-5
Goldhammer, F., & Zehner, F. (2017). What to make of and how to interpret process data. Measurement: Interdisciplinary Research and Perspectives, 15(3–4), 128–132. https://doi.org/10.1080/15366367.2017.1411651
He, S. and Cui,Y. (2025). A systematic review of the use of log-based process data in computer-based assessments. Comput. Educ. 228, C (Apr 2025). https://doi.org/10.1016/j.compedu.2025.105245
Availability and Gaps in Educational Research Data in a National Social Science Repository
Data sharing is increasingly recognised as essential for transparency, reproducibility, and reuse in the social sciences (Heers, 2023). National repositories play a key role in supporting researchers to share high-quality, FAIR data while providing professional curation, controlled access, and documentation (Late & Ochsner, 2024). SWISSUbase, Switzerland’s main social science data repository, exemplifies such infrastructure, enabling researchers across disciplines to deposit and disseminate datasets.
This contribution zooms in on educational research within SWISSUbase. Using repository metadata, we profile datasets originating from education studies, including data types, institutional affiliations, and linguistic regions. These patterns are compared with publicly available information on education research outputs funded by the Swiss National Science Foundation, revealing areas where educational research is underrepresented or absent from the repository.
The analysis will highlight structural and cultural barriers that may limit data sharing in education research and relate to those found in previous research (van der Zee & Reich, 2018). The discipline lags somewhat behind others in sharing and depositing data (Logan, Hart, & Schatschneider). We discuss strategies to better engage education researchers, including tailored support, guidance on ethical and practical aspects of data sharing, and incentives to deposit. By highlighting these gaps and possible interventions, the contribution aims to foster debate on how repositories can support a more comprehensive, FAIR-aligned education research data landscape.
References:
Heers, M. (2023). Data sharing in the Social Sciences. FORS Guides, 21, Version 1.1, 1-10. https://doi.org/10.24449/FG-2023-00021
Late, E. & Ochsner, M. (2024). Re-use of research data in the social sciences. Use and users of digital data archive. PLoS ONE 19(5): e0303190. https://doi.org/10.1371/journal.pone.0303190
Logan, J. A. R., Hart, S. A., & Schatschneider, C. (2021). Data Sharing in Education Science. AERA Open, 7, 23328584211006475. https://doi.org/10.1177/23328584211006475
van der Zee, T., & Reich, J. (2018). Open Education Science. AERA Open, 4(3), 2332858418787466. https://doi.org/10.1177/2332858418787466
| ||
