Skip to content
le chonk
Menu

Le Chonk benchmarks

A number is useful when you know where it came from. These are reported preview results, with source links and the conditions we have not verified.

On this page

What has actually been checked?

We checked Mistral’s release announcement for the results below. The announcement attributes these coding results to Artificial Analysis. We have not directly verified the external evaluator’s run pages or reproduced the tests.

Evidence status: vendor-reported, attributed to Artificial Analysis. This site has no original Le Chonk benchmark measurements. A source being official does not make its performance claims independent.

These results refer to Mistral Large 4 public preview. A later API update can change performance without keeping the original preview fixed.

The reported numbers

Mistral Large 4 public preview · vendor-reported results
BenchmarkResultEvidenceWhat is missing
DeepSWEv1.161.7%Vendor-reportedMistral → Artificial Analysis ↗Mistral Large 4 public previewAgent harness, inference budget and repeated-run uncertainty not independently checked.
SWE-Atlas-QnANot specified in release59.4%Vendor-reportedMistral → Artificial Analysis ↗Mistral Large 4 public previewDataset revision and evaluation configuration not independently checked.
Terminal-Bench4.028.3%Vendor-reportedMistral → Artificial Analysis ↗Mistral Large 4 public previewAgent harness, tool permissions and inference budget not independently checked.
Coding Agent IndexNot specified in release49.8%Vendor-reportedMistral → Artificial Analysis ↗Mistral Large 4 public previewComposite result, not a separate task pass rate. Weighting not independently checked.

The rows stay in topic order. They are different tests, not a cross-benchmark ranking. Missing evaluator settings are left as missing.

What do these scores measure?

Software engineering tasks

For repository work, look beyond a pass percentage. Ask which tasks were attempted, what files and tools were available, and how success was judged. A useful patch also needs to be understandable and avoid breaking unrelated behavior.

Terminal tasks

A terminal evaluation depends on the environment, commands the agent can run, time limits and the harness around the model. An agent that gets a larger budget may complete work another agent cannot finish in time. Keep those conditions attached to the score.

Composite indexes

A combined index compresses several measurements into one number. Before using it to choose a model, check the underlying task mix and weighting. Your application may care about a small part of the index or about behavior it does not measure at all.

Why we do not name an overall winner

We do not have a complete, directly checked comparison table with matching model snapshots, budgets and environments. Publishing a ranked list would hide those gaps. The reported results are useful leads for further testing, not a replacement for it.

Conditions to check before comparing

Evidence audit · all rows refer to the reported preview results above
ConditionCurrent statusWhy it matters
Exact model artifactPublic preview named; immutable revision not verifiedA rolling endpoint can change between runs.
Dataset revisionDeepSWE v1.1 and Terminal-Bench 4 stated; others incompleteDifferent revisions can contain different tasks.
Agent harness and toolsNot independently verifiedThe same model can behave differently with different tools.
Reasoning and token budgetNot independently verifiedMore attempts or longer reasoning can change results and cost.
Sampling and repeated runsNot independently verifiedA single reported value does not show run-to-run variation.
Independent evaluation pageNot directly checkedAttribution in a release is not direct verification.
This site’s reproductionsNoneDo not treat these numbers as our own experiments.

Do not compare numbers from different dataset versions as if they were the same race. If the comparison source does not disclose a condition, record that gap instead of filling it with an assumption.

What a site-run test would include

No original test results are available yet. When we add them, they will sit in their own section with the exact model, date, prompts, harness version, settings and sample size. API tests and local weight tests will be labeled separately.

  1. Choose tasks and success criteria before running the models.
  2. Pin model versions where possible; record rolling endpoints as rolling.
  3. Use the same tools, input files and budget for each model.
  4. Save failures and retries as well as successful outputs.
  5. Report accepted-task cost and latency alongside scores.
  6. Publish limitations and enough material to repeat the test.

A small, honest task set can help answer a specific question. It should not be presented as proof of general superiority.

Download the evidence table

The files include metric versions, the preview model name, reported values, source URLs, dates, attribution and missing conditions. The downloadable data uses the same source file as this page.

Sources