Automated scoring systems for open-ended questions in Dutch education

In Dutch secondary education, open-ended exam questions (such as writing an email or short essay) are still graded manually for the course of Dutch. This process is not only time-consuming for teachers but can also lead to inconsistencies. Even with standardized rubrics, scoring can vary depending on how individual graders interpret student responses. This project explored whether AI could support teachers by providing reliable, rubric-aligned scores for the aspects of content, language, and presentation, and to do so in a way that is transparent and explainable.

To investigate this, this project compared two approaches. The first was fine-tuning Dutch-specific transformer models such as RobBERT. RobBERT performed best on the language aspect, detecting and evaluating grammar and spelling, while content and presentation proved more difficult to model due to rubric variability and limited data. The second approach relied on GPT-4 in a prompt-based setup. GPT-4 achieved competitive results, especially for content and language, with few-shot prompting and justification requests further improving its alignment with the rubrics. Both approaches had strengths and limitations, which made their comparison especially valuable.

Just as importantly, this project focused on explainability. For the fine-tuned models, I used SHapley Additive Explanations  (SHAP) analysis to highlight which words and features most strongly influenced predictions. For GPT-4, we designed prompts that asked the model to provide written justifications for its scores. These explainability techniques could help examiners see why a score was assigned, increasing trust in the system.

A central insight was that the challenges in automated scoring are not only technical but also human. Expert reviews revealed that even trained graders sometimes disagreed, and the data itself contained inconsistencies. This raised an important question: if the “gold standard” is not always reliable, how should we evaluate and train AI systems to support fair grading? This underlined both the potential of AI to assist in creating more consistent scoring and the need for further research into how automated systems and human graders can best complement each other.

Overall, the results suggest that a hybrid system, combining the precision of fine-tuned models with the flexibility of generative large language models, could be especially effective in supporting fair and efficient grading in Dutch education.

Student

  • Nafsika Lachana

Academic supervisor(s)

  • Dr. Matthieu Brinkhuis
  • Lientje Maas (CITO)

Grant funding agency and (co-)funding non-academic partners