An independent, scientifically rigorous evaluation of AI beyond English
Understand how frontier models perform across languages and enterprise agentic tasks, cultural contexts, and multimodal domains to choose the best model for every market you serve.
Standard benchmarks mask frontier models' performance regressions when used in any language but English. AURORA stress-tests LLMs under realistic, localized constraints using natively-authored tasks to map the true boundaries of global AI capability.
Task pass rate across 10 languages (5 trials per task).
Mean pass rates over 324 tasks (5 trials per task). Whiskers show 95% Wilson score intervals.
Pass rate against release date, to show whether a later model is actually a better one.
Each dot is one model at its vendor release date; whiskers show 95% Wilson score intervals. Left to right is time, so a dot that sits high and early led the field when it shipped.
*In Terminal-Bench, the GLM 5.3 Flash scores are on a quantized model as this is pinned by the service provider and cannot be changed
Agent reliability per model on the airline, retail, telecom and banking domains, in English, German, Korean, Turkish and Persian.
pass^1 is the probability of passing a single trial.
gemini-3.5-flash, temperature 1.0, with the system prompt tuned for realistic behaviorgpt-5.6-sol, claude-opus-5, CohereLabs/command-a-plus-05-2026, gemini-3.8-flashnone, Opus 5 off (banking: low), Command A+ has no off switch, Gemini 3.8 Flash low (cannot be disabled). Temperature: provider default.gpt-4.1-2025-04-14 (τ³-bench default) for airline and telecom, deepseek-v4-flash for retail and bankingbase split; banking: stratified 20 task sample, same task ids in every locale and for every agentMulti-turn conversational accuracy per model by language.
Some languages take more tokens than English, and are thus more expensive to run. For example, the same dialogue costs about 1.4× as much in Korean or Arabic as in English.
You cannot measure multilingual performance by machine-translating an English benchmark. Doing so introduces a non-native text distribution, does not result in culturally relevant tasks, and may introduce factual errors.
We compare two multilingual versions of GAIA: the MAPS-GAIA machine translations of the original English tasks, and our carefully audited and culturally adapted tasks.
Models score an average of 20.7 points higher on the audited, culturally adapted benchmark, reflecting measurement error introduced by translation in the original multilingual benchmark.
Pass rate on the MAPS-GAIA machine-translated benchmark compared to LILT's audited GAIA-v2-LILT.
165 query–answer pairs per language · the dashed line marks each model's English score
Table 2 · Source: arXiv:2604.24929
Share of tasks the audit revised at all.
Table 1 · Source: arXiv:2604.24929
Word- and character-level edit distance from the machine translation.
Table 1 · Source: arXiv:2604.24929
pass rate (equivalently pass@1 or pass^1) is the chance a model will successfully pass a single task. We estimate it by having each model complete each task multiple times.
pass@k is the chance that a model will pass a task at least once out of k attempts. Intuitively, this is best used for difficult tasks that could reasonably be solved by running the model multiple times and selecting the best answer. It is estimated using the unbiased estimator (see Appendix A of Chen et al., 2021).
pass^k is the chance that a model will pass a task reliably in all k attempts. This metric is best used for tasks that can't be rerun (like interactions with a customer, as in τ³-bench) or where failure is expensive or dangerous. It is estimated using the unbiased estimator introduced in the first τ-bench paper (Yao et al., 2024).
LILTBench is an open collection of 92 coding tasks across 31 languages, authored by community teams to break frontier models in their own language. Tasks ship with the original-language instructions and an English translation, so the same problem can be run either way to see how much prompt language alone moves performance.
The tasks, the harness, and the leaderboard are public.