Agentic Understanding and Reasoning in Original-language, Real-world Assessments

AURORA: LILT Multilingual
AI Leaderboard

An independent, scientifically rigorous evaluation of AI beyond English

Understand how frontier models perform across languages and enterprise agentic tasks, cultural contexts, and multimodal domains to choose the best model for every market you serve.

Standard benchmarks mask frontier models' performance regressions when used in any language but English. AURORA stress-tests LLMs under realistic, localized constraints using natively-authored tasks to map the true boundaries of global AI capability.

Public Benchmark Results

Ranked by pass rate

Task pass rate across 10 languages (5 trials per task).

Language

Mean pass rates over 324 tasks (5 trials per task). Whiskers show 95% Wilson score intervals.

By release date

Pass rate against release date, to show whether a later model is actually a better one.

Each dot is one model at its vendor release date; whiskers show 95% Wilson score intervals. Left to right is time, so a dot that sits high and early led the field when it shipped.

How these models were run
Evaluation framework
Harbor v0.21.0
Agent harness
Terminus-2 (following the original Terminal-Bench)
Trials per task
5
Time limits
Per-task, adjusted based on task complexity by each task author, minimum 600s
Task instructions
Authored by native speakers in each language
Reasoning
Provider defaults, unless otherwise noted

*In Terminal-Bench, the GLM 5.3 Flash scores are on a quantized model as this is pinned by the service provider and cannot be changed

Multilingual τ³-benchMulti-turn tool use in customer support scenarios

Agent reliability

Agent reliability per model on the airline, retail, telecom and banking domains, in English, German, Korean, Turkish and Persian.

Metric
Group by
Domain
Claude Opus 5.5GPT-6-SolClaude Opus 5GPT-5.6-SolGemini 3.8 FlashCommand A+

pass^1 is the probability of passing a single trial.

How these models were run
User simulator
gemini-3.5-flash, temperature 1.0, with the system prompt tuned for realistic behavior
Agent models
gpt-5.6-sol, claude-opus-5, CohereLabs/command-a-plus-05-2026, gemini-3.8-flash
Agent reasoning
Disabled where the model allows it, else lowest effort, to simulate low-latency customer support: Sol none, Opus 5 off (banking: low), Command A+ has no off switch, Gemini 3.8 Flash low (cannot be disabled). Temperature: provider default.
Serving
OpenRouter for all agents except Command A+, which ran on a local vLLM deployment
Assertion judge
gpt-4.1-2025-04-14 (τ³-bench default) for airline and telecom, deepseek-v4-flash for retail and banking
Task split
Airline, telecom, retail: base split; banking: stratified 20 task sample, same task ids in every locale and for every agent
Modifications
Thorough localization of user simulator prompts, task scenarios and database contents; minor edit to agent system prompt instructing it to use a specific language
Trials per task
4
Agent scaffold
official τ³-bench scaffold, modified to support multilinguality
Max tokens per turn
4096
Multilingual MultiChallenge ↗Long-context instruction-following

Multi-turn accuracy (pass@1)

Multi-turn conversational accuracy per model by language.

Group by
GPT-5.6-SolClaude Opus 5.5Gemini 3.8 FlashClaude Opus 5GPT-6-SolClaude Opus 4.6Gemini 3.1 Flash-LiteGPT-5.2Kimi-K3Command-A+

Tokens per dialogue

Some languages take more tokens than English, and are thus more expensive to run. For example, the same dialogue costs about 1.4× as much in Korean or Arabic as in English.

Scale
Multilingual GAIA-v2 · LILT ↗Agentic reasoning and tool use

You cannot measure multilingual performance by machine-translating an English benchmark. Doing so introduces a non-native text distribution, does not result in culturally relevant tasks, and may introduce factual errors.

We compare two multilingual versions of GAIA: the MAPS-GAIA machine translations of the original English tasks, and our carefully audited and culturally adapted tasks.

Models score an average of 20.7 points higher on the audited, culturally adapted benchmark, reflecting measurement error introduced by translation in the original multilingual benchmark.

pass@1 accuracy

Pass rate on the MAPS-GAIA machine-translated benchmark compared to LILT's audited GAIA-v2-LILT.

Group by

165 query–answer pairs per language · the dashed line marks each model's English score

Table 2 · Source: arXiv:2604.24929

% tasks revised

Share of tasks the audit revised at all.

Table 1 · Source: arXiv:2604.24929

Word & char edit rates

Word- and character-level edit distance from the machine translation.

Table 1 · Source: arXiv:2604.24929

How this benchmark was built and run
Tasks
165 query–answer pairs per language, 825 in total across the five languages
Languages
Arabic, German, Hindi, Korean, Portuguese (BR), plus original English
Scoring
Follows the English benchmark, with modifications to handle commas as decimal points and other locale-specific formatting
Max turns
12 steps for the manager, 20 for the search subagent, following the English benchmark
Human review
Each task was adapted by bilingual annotators, then reviewed by a LILT researcher

Metrics Definitions

01 pass rate

pass rate (equivalently pass@1 or pass^1) is the chance a model will successfully pass a single task. We estimate it by having each model complete each task multiple times.

02 pass@k

pass@k is the chance that a model will pass a task at least once out of k attempts. Intuitively, this is best used for difficult tasks that could reasonably be solved by running the model multiple times and selecting the best answer. It is estimated using the unbiased estimator (see Appendix A of Chen et al., 2021).

03 pass^k

pass^k is the chance that a model will pass a task reliably in all k attempts. This metric is best used for tasks that can't be rerun (like interactions with a customer, as in τ³-bench) or where failure is expensive or dangerous. It is estimated using the unbiased estimator introduced in the first τ-bench paper (Yao et al., 2024).

LILTBench

LILTBench is an open collection of 92 coding tasks across 31 languages, authored by community teams to break frontier models in their own language. Tasks ship with the original-language instructions and an English translation, so the same problem can be run either way to see how much prompt language alone moves performance.

The tasks, the harness, and the leaderboard are public.