Conference Agenda
Overview and details of the sessions of this conference. Please select a date or location to show only sessions at that day or location. Please select a single session for detailed view (with abstracts and downloads if available).
Please note that all times are shown in the time zone of the conference. The current conference time is: 15th Sept 2026, 09:35:16am EEST
External resources will be made available 5 min before a session starts. You may have to reload the page to access the resources.
|
Daily Overview |
| Session | |
|
STE-R PS1: Remote Session 1 Location: online Session Chair: Johannes Kubasch, University of Wuppertal | |
| Presentation 5 | |
3:42pm - 4:00pm
An Open-Source Benchmark for Enhancing Image Interpretation Across Four Scientific Disciplines Using Multimodal Large Language Models (MLLMs) in Low-Resource Environments Grhapes, INSEI / CY Cergy Paris Université, France Abstract. This article presents a new framework designed to improve image interpretation in four high school science disciplines through the use of open-weight language models optimized for resource-constrained environments. It also makes a significant contribution to the field of educational AI. While most previous studies focus on proprietary or hybrid models, evaluated using benchmarks such as VisioMath, ScienceQA, SceMQA, or MMMU, our study proposes an alternative approach. We optimized four open-weight models—Mistral-7B, LLaVA-1.5-7B, Kosmos-2, and Qwen2-VL—selected for their lightweight architecture, low memory footprint, and robustness. The optimization focused on low-rank adaptation (LoRA) and 4-bit quantization in a resourceconstrained environment, specifically to improve the understanding of scientific images. The results of Bench Low, distinguishing between multiple-choice (MC) and free-response (FR) questions, reveal significant variations depending on the type of task, model, and discipline. LLaVA-1.5-7B stood out in particular for its superior performance, achieving near-perfect accuracy on multiple-choice questions (96.9%) and the highest score on open-ended questions (69.1%), with an overall average of 79.3%. In contrast, Kosmos-2 achieved the lowest results (40.8% overall), with high variability between disciplines, reflecting uneven coverage of training data. For reproducibility purposes, the source codes are fully available on Github: https://github.com/mouazmikail/Bench-Lowv1/tree/master | |
