Can large language models perform as reliably in Hungarian or Italian as they do in English? To help answer that question, DG Translation has released the EU MMLU, a high-quality multilingual benchmarking dataset designed to assess whether LLMs perform fairly and effectively across the EU’s linguistic diversity.
Most AI benchmarking datasets are built in English, which means a model may perform very well in English while underperforming in other languages. The EU MMLU helps address this gap by building on one of the datasets most widely used to evaluate LLMs: Massive Multitask Language Understanding (MMLU). This dataset contains thousands of multiple-choice questions covering 57 subjects, from science to law.
The EU MMLU focuses on 7 of those subject areas, selected for their relevance to the EU. To ensure that each question retains the same meaning and level of difficulty across languages, 1000+ benchmark questions were translated and revised with the help of nearly 240 student translators and project managers from 21 universities in the European Master’s in Translation (EMT) network.
The EU MMLU is currently available in 16 EU official languages, with more to come.
Alongside the dataset, DG Translation has also published a list of core quality criteria for EU-oriented multilingual benchmarking, grounded in European institutions, legislation and realities. Together, the EU MMLU dataset and the list of core quality criteria are designed to improve how today’s language models are evaluated – and to help shape the next generation of multilingual AI.