On this page
What has actually been checked?
We checked Mistral’s release announcement for the results below. The announcement attributes these coding results to Artificial Analysis. We have not directly verified the external evaluator’s run pages or reproduced the tests.
Evidence status: vendor-reported, attributed to Artificial Analysis. This site has no original Le Chonk benchmark measurements. A source being official does not make its performance claims independent.
These results refer to Mistral Large 4 public preview. A later API update can change performance without keeping the original preview fixed.
The reported numbers
| Benchmark | Result | Evidence | What is missing |
|---|---|---|---|
| DeepSWEv1.1 | 61.7% | Vendor-reportedMistral → Artificial Analysis ↗Mistral Large 4 public preview | Agent harness, inference budget and repeated-run uncertainty not independently checked. |
| SWE-Atlas-QnANot specified in release | 59.4% | Vendor-reportedMistral → Artificial Analysis ↗Mistral Large 4 public preview | Dataset revision and evaluation configuration not independently checked. |
| Terminal-Bench4.0 | 28.3% | Vendor-reportedMistral → Artificial Analysis ↗Mistral Large 4 public preview | Agent harness, tool permissions and inference budget not independently checked. |
| Coding Agent IndexNot specified in release | 49.8% | Vendor-reportedMistral → Artificial Analysis ↗Mistral Large 4 public preview | Composite result, not a separate task pass rate. Weighting not independently checked. |
The rows stay in topic order. They are different tests, not a cross-benchmark ranking. Missing evaluator settings are left as missing.
What do these scores measure?
Software engineering tasks
For repository work, look beyond a pass percentage. Ask which tasks were attempted, what files and tools were available, and how success was judged. A useful patch also needs to be understandable and avoid breaking unrelated behavior.
Terminal tasks
A terminal evaluation depends on the environment, commands the agent can run, time limits and the harness around the model. An agent that gets a larger budget may complete work another agent cannot finish in time. Keep those conditions attached to the score.
Composite indexes
A combined index compresses several measurements into one number. Before using it to choose a model, check the underlying task mix and weighting. Your application may care about a small part of the index or about behavior it does not measure at all.
Why we do not name an overall winner
We do not have a complete, directly checked comparison table with matching model snapshots, budgets and environments. Publishing a ranked list would hide those gaps. The reported results are useful leads for further testing, not a replacement for it.
Conditions to check before comparing
| Condition | Current status | Why it matters |
|---|---|---|
| Exact model artifact | Public preview named; immutable revision not verified | A rolling endpoint can change between runs. |
| Dataset revision | DeepSWE v1.1 and Terminal-Bench 4 stated; others incomplete | Different revisions can contain different tasks. |
| Agent harness and tools | Not independently verified | The same model can behave differently with different tools. |
| Reasoning and token budget | Not independently verified | More attempts or longer reasoning can change results and cost. |
| Sampling and repeated runs | Not independently verified | A single reported value does not show run-to-run variation. |
| Independent evaluation page | Not directly checked | Attribution in a release is not direct verification. |
| This site’s reproductions | None | Do not treat these numbers as our own experiments. |
Do not compare numbers from different dataset versions as if they were the same race. If the comparison source does not disclose a condition, record that gap instead of filling it with an assumption.
What a site-run test would include
No original test results are available yet. When we add them, they will sit in their own section with the exact model, date, prompts, harness version, settings and sample size. API tests and local weight tests will be labeled separately.
- Choose tasks and success criteria before running the models.
- Pin model versions where possible; record rolling endpoints as rolling.
- Use the same tools, input files and budget for each model.
- Save failures and retries as well as successful outputs.
- Report accepted-task cost and latency alongside scores.
- Publish limitations and enough material to repeat the test.
A small, honest task set can help answer a specific question. It should not be presented as proof of general superiority.
Download the evidence table
The files include metric versions, the preview model name, reported values, source URLs, dates, attribution and missing conditions. The downloadable data uses the same source file as this page.
Sources
- Mistral Large 4 release announcement
API preview, planned weight release, 49B active parameters and reported benchmark results.
- Mistral Large model documentation
1.05T total, 52B active, 1M context. Displays both $0.68 / $2.09 and higher $1.36 / $4.18 prices. No discount end date confirmed.