Among the models currently evaluated on EuroEval’s Dutch leaderboard, Gemini 3.7 Flash ranks first as of September 2026, ahead of Gemini 3.6 Flash and gpt-5.6-sol. That does not make it the newest option: Gemini 3.8 Flash arrived in September and also GPT-6.1 Sol, and neither has been benchmarked in Dutch yet.
The ranking is also a poor way to pick a model. There is no single best AI model for Dutch, because being “good at Dutch” is not one skill but at least three: processing the language, knowing the country, and writing for a Dutch reader.
The short version
- Among evaluated models, Gemini 3.7 Flash leads EuroEval’s Dutch board with a rank score of 1.24, where lower is better.
- Dutch language handling is settled. The seventeen models in EuroEval’s top five ranks sit within 2.5 points of each other on Dutch sentiment.
- Knowledge of the Netherlands is not settled. Those same seventeen models range from 22.1 to 84.2, so the model you pick decides the outcome.
Below are the per-model scores, what the benchmarks do and do not measure, and how to test which model handles the Dutch you actually work with.
Which is the best AI model for Dutch right now?
Gemini 3.7 Flash leads EuroEval’s Dutch leaderboard with a rank score of 1.24. EuroEval, known until last year as ScandEval, tests models on nine Dutch datasets and averages their position across those nine, so a lower score is better. Gemini 3.6 Flash follows at 1.34 and gpt-5.6-sol at 1.43. The board listed 786 models when we read it.
The first surprise is which models come out on top. They are the fast, cheap Flash variants rather than the largest flagships. For Dutch-language work that is good news: you do not need the most expensive model available.
| Model (as EuroEval lists it) | Rank | Sentiment (DBRD) | Dutch local knowledge (MultiLoKo-nl) | Plain-language rewriting | Open weights |
|---|---|---|---|---|---|
| gemini-3.7-flash | 1 | 92.0 | 83.0 | 49.2 | No |
| gemini-3.6-flash | 2 | 91.5 | 81.3 | 47.9 | No |
| gpt-5.6-sol | 3 | 93.6 | 67.5 | 54.0 | No |
| claude-sonnet-4-5 (thinking) | 3 | 93.4 | 33.8 | 49.7 | No |
| Qwen3.6-27B-FP8 | 3 | 91.6 | 24.3 | 50.5 | Yes |
| gemma-4-31B-it | 4 | 93.4 | 28.5 | 53.4 | Yes |
| grok-4-1-fast-reasoning | 4 | 93.6 | 48.0 | 50.8 | No |
| EuroLLM-9B-Instruct-2512 | 13 | 92.1 | 21.7 | 56.6 | Yes |
All scores come from EuroEval, in September 2026, rounded to one decimal. All eight models were run on the validation split, so you can compare the rows with each other.
Why do the best models look almost identical on Dutch?
On ordinary Dutch language tasks the gaps between the leading models are too small to mean anything. DBRD shows this well. It scores sentiment classification on Dutch book reviews from the review site Hebban.nl, and across the seventeen models in EuroEval’s top five ranks the scores run from 91.5 to 94.0. That is a 2.5-point spread on a scale of 100, between models that differ enormously in size and price.
Reading comprehension tells the same story. On SQuAD-nl the same group scores between 70.8 and 81.0, another narrow band.
Among leading general-purpose models, ordinary Dutch performance is tightly clustered. Pick on language ability alone and you are picking at random.
Where do the differences actually show up?
The real separation happens on knowledge of the Netherlands. MultiLoKo-nl tests exactly that, and its Dutch questions were written natively rather than translated, targeting topics that matter locally. The underlying paper describes MultiLoKo as a local-knowledge benchmark spanning 31 languages. One sample question asks when the television journalist Twan Huys and the then prime minister Mark Rutte staged a Dutch edition of the Correspondents’ Dinner at the Beurs van Berlage, the old Amsterdam exchange building.
Across those same seventeen models in the top five ranks, scores on this test run from 22.1 to 84.2. Gemini 3.7 Flash reaches 83.0, gpt-5.6-sol 67.5, Claude Sonnet 4.5 with thinking 33.8 and Qwen3.6-27B-FP8 24.3. These are the same models that sit within two points of each other on sentiment.
The second weak spot hits Dutch organisations directly. The Duidelijke Taal dataset measures whether a model can rewrite a complicated Dutch sentence into plain language. The source material comes from the Instituut voor de Nederlandse Taal, the Dutch Language Institute. It holds 6,986 original and 6,986 simplified sentences, each rated by crowdworkers for simplicity, accuracy and fluency. The best score in the top five ranks is 55.6, lower than the best score those models reach on any other Dutch task in the set.
That is precisely the job Dutch municipalities, care providers and insurers want AI to do. The Dutch government writes its own web copy at B1 level, the grade its own guidance says most adult readers handle comfortably. CBS, the Dutch statistics office, reports that three million adults aged 16 to 75 struggle with reading, writing or arithmetic. Plain-language rewriting is not a nice-to-have, and it is the Dutch task these models handle worst.
Can you trust a Dutch benchmark?
You can trust a Dutch benchmark in part, and you should know which part. Two of EuroEval’s nine Dutch datasets are translations from English. Google Translate produced SQuAD-nl, and eight students then corrected the test set by hand. Winogrande-nl is a translated and filtered version of the English original. The other seven use real Dutch material: book reviews from Hebban.nl, Dutch WikiHow articles, sentences from the SoNaR corpus.
But a dataset written in Dutch is not automatically a dataset about the Netherlands. CoNLL-nl consists of annotations from the Belgian newspaper De Morgen, dating from 2000. INCLUDE-nl draws on Flemish exam material, including sentences about student housing that use Flemish words rather than the words used in the Netherlands. Both datasets are useful, but neither is the Dutch a municipality in Groningen writes. And only one dataset, MultiLoKo-nl, was written to test what a model knows about Dutch life itself.
The older Dutch benchmark DUMB, from the University of Groningen, still turns up in these discussions but measures something else entirely. DUMB compares fourteen encoder models such as DeBERTaV3 and XLM-R. It says nothing about ChatGPT, Claude or Gemini.
EuroEval flags one more caveat itself. The top entries were run on the validation split, and its FAQ says those scores “aren’t strictly comparable to the regular test-split scores”. Small gaps at the top are noise. So treat any leaderboard as a starting point and not a verdict, which is the same lesson that applies to LLM evaluation generally.
What do OpenAI, Google and Anthropic publish about Dutch?
All three evaluate their models in many languages, but almost none of them publish a Dutch number you can read off. OpenAI’s multilingual benchmark MMMLU covers exactly fourteen languages: Arabic, Bengali, German, Spanish, French, Hindi, Indonesian, Italian, Japanese, Korean, Portuguese, Swahili, Yoruba and Chinese. Dutch is not among them.
Anthropic’s multilingual documentation uses that same fourteen-language set, so it carries no Dutch row. Its system card for Claude Sonnet 4.6 goes wider. That document runs Global MMLU across 42 languages and puts Dutch in the high-resource group, but it reports one average of 91.0 percent for that whole group rather than a Dutch score. Anthropic has also published research on how Claude’s values shift by language, covering the twenty most common languages on Claude.ai, which finds that Claude “leans toward candor in Dutch”.
Google evaluates Gemini 3 Pro on MMMLU and Global PIQA without breaking Dutch out. Mistral lists Dutch in its language overview, in the group of languages with “strong expected performance”, but publishes no figures for it.
So Dutch is starting to appear inside vendor evaluations, just almost never as a figure. Dutch does have its own rows in Cohere Labs’ Global-MMLU and in Meta’s Belebele. If you want a number, independent benchmarks are still where it lives.
Is a Dutch or European model the safer choice?
On quality, a Dutch or European model is not the better choice yet. A City of Amsterdam study published on 10 August 2026 reached that finding. Its authors assessed more than thirty models on factuality, honesty, bias, energy use, cost and training-data transparency. On Dutch factuality GPT-5 scores highest at 0.76. The Dutch fine-tunes trail badly: GEITje 7B Ultra reaches 0.39 and Fietje 2 reaches 0.43. EuroLLM 9B scores 0.44 and EuroLLM 22B 0.47.
The same study supplies the best argument against ranking models on a single number. HonestCityBench, a purpose-built Dutch benchmark of 530 prompts, measures honesty. GPT-4o leads there at 0.43, while GPT-5 manages only 0.14. The authors sum it up this way: “factuality and honesty are governed by distinct properties, with high factuality not implying high honesty”. One limit is worth noting. The study only covered models available through the city’s own Azure environment, so Claude is absent from it.
You cannot buy the Dutch option yet either. GPT-NL is a collaboration between the research organisation TNO, the Netherlands Forensic Institute and the IT cooperative SURF, backed by 13.5 million euros from the Netherlands Enterprise Agency (RVO). It is training a 26-billion-parameter model on rights-cleared Dutch data. Its HuggingFace organisation is still empty, and the model is open only to GPT-NL’s launching customers. Its own definition of success is candid about the ambition. The document aims for performance “comparable to the Llama2 7B model and GPT-3 175B (or equivalent) models”.
One number runs against that pattern. On plain-language rewriting, EuroLLM-9B-Instruct scores 56.6, higher than every commercial flagship in our table. European models are not uniformly weaker. They are weak at different things. If open weights are on your shortlist, our guide to open-source LLMs covers the wider trade-off.
One more lesson: GEITje, the best-known Dutch model, lost its original weights. Its model page says all weights and checkpoints were deleted at the pressing request of Stichting BREIN, the Dutch copyright enforcement body. In his takedown post, its creator Edwin Rijgersberg explains that the Dutch training corpus contained copyrighted material.
How much more does Dutch cost in tokens?
Dutch costs more tokens than English for the same content. That gap is smaller than it was a few years ago. We measured it on the Europarl corpus using 4,747 parallel sentence pairs. With the old GPT-4 tokenizer (cl100k_base) the Dutch text needed 56 percent more tokens than the English. With o200k_base, the tokenizer behind every current OpenAI model, the penalty is down to 16 percent.
That 16 percent costs you twice. You pay per token, and your context window fills up faster. Long compound nouns are the worst offenders. “Arbeidsongeschiktheidsverzekering” costs eight tokens against three for “disability insurance”.
A shortage of Dutch text is not the problem. In the most recent Common Crawl snapshot (CC-MAIN-2026-34), 1.78 percent of pages are Dutch, ranking it eleventh of 161 languages. There is plenty of Dutch on the web. What is scarce is Dutch evaluation.
How do you choose a model for Dutch work?
You should start from the task, not from the leaderboard. This order works in practice:
- Decide which of the three skills you need. For processing language, almost any current model will do and you should choose on price and latency. For Dutch facts, people or regulations, the model you choose decides the outcome.
- Never rely on the model to hold knowledge of the Netherlands. That 62-point gap on MultiLoKo-nl mostly disappears when you supply the facts yourself, through RAG or a fixed source. Do not trust what the model memorised.
- Test on your own material. Take fifty real documents from your organisation, with your own jargon and abbreviations, and give two or three models the same task.
- Keep a human in the loop for plain-language work. Scores here are the weakest of any Dutch task, so you still need an editing round.
- Only then weigh the constraints: data residency, cost per token, and what the EU AI Act has required since August 2026.
Frequently asked questions (FAQ)
Which AI model is best at Dutch?
Among the models EuroEval has evaluated, Gemini 3.7 Flash led its Dutch board with a rank score of 1.24. Newer releases such as Gemini 3.8 Flash and GPT-6 Astra are not on that board yet. For pure language work the top models are nearly interchangeable; for Dutch knowledge the gap is wide.
Is ChatGPT good at Dutch?
ChatGPT handles Dutch well as a language. On Dutch sentiment classification gpt-5.6-sol scores 93.6 out of 100, matching the other leading models. On knowledge of specifically Dutch topics it reaches 67.5, while the Gemini Flash models clear 81. For factual questions about the Netherlands that difference is noticeable in practice.
Is there a good Dutch-built language model?
No Dutch-built model matches the international models yet. GPT-NL, built by TNO, the Netherlands Forensic Institute and SURF, is still training and was open to launching. The Dutch fine-tunes GEITje 7B Ultra and Fietje 2 score lower on Dutch factuality than large general models in the City of Amsterdam study.
Why is there no Dutch score in OpenAI’s and Anthropic’s benchmarks?
OpenAI’s multilingual benchmark MMMLU covers fourteen languages and Dutch is not one of them. Anthropic goes wider: its Claude Sonnet 4.6 system card runs Global MMLU across 42 languages with Dutch in the high-resource group, but reports only that group’s average rather than a Dutch figure. For Dutch numbers you still need independent benchmarks such as EuroEval.
Does Dutch cost more than English on an AI model?
Dutch costs roughly 16 percent more tokens than English for the same content with the current OpenAI tokenizer, measured on a parallel corpus of 4,747 sentence pairs. With the older GPT-4 tokenizer the penalty was 56 percent. Since you pay per token, that difference shows up in your bill and in how fast your context window fills.
