Automatic annotation of Dutch educational assessment questions using Large Language Models

This project focuses on the automatic evaluation of curriculum alignment. In simple terms, curriculum alignment refers to the extent to which learning objectives, teaching activities, and assessments are coherently connected. Traditionally, measuring this alignment is a time-consuming and often subjective process, because it involves checking all educational materials against the intended learning goals. This project explores how artificial intelligence could help make this process faster, more objective, and scalable.
To address this, the research explores the use of large language models (LLMs) to automate the annotation of Dutch assessment questions with subject-specific concepts. Specifically, it investigates two different types of models (GPT-4.1 nano and mBERT models) using a labeled dataset of Dutch statistics questions.
The results were promising. Both models showed strong potential, with mBERT performing particularly well, achieving an accuracy of 91.7%. I also observed that the performance of the models varied depending on the type of questions and the subject matter, highlighting the importance of careful adaptation and evaluation when applying these models for different educational contexts.
By automating question annotation, this research can help educators better align assessments with learning objectives, ultimately improving the learning experience for students. More broadly, it contributes to the growing integration of AI in education, offering insights into which approaches are best suited for different scenarios and demonstrating how technology can support teachers and curriculum designers in their work.
Student
- Isabela Motoki
Academic supervisor(s)
- Dr. Matthieu Brinkhuis
- Lientje Maas (CITO)