Four labels, four different meanings
| Label | Meaning |
|---|---|
| Verified documentation | We checked a primary source for this specification or price. This does not mean we tested the model. |
| Vendor-reported | The model provider reported the result, possibly citing an outside evaluator. We have not independently reproduced it. |
| Estimated | The number comes from a stated formula and inputs, not a hardware or API measurement. |
| Pending / Not verified | A release is planned or the evidence needed to confirm the claim is missing. Unknown numbers are stored as null, never zero. |
Source rules
We prefer model documentation, provider pricing pages, official repositories and original evaluation records. Search snippets help find sources; they are not a substitute for reading the source page.
Each important data record carries a source URL, check date, model version, evidence type and notes. Publication dates are stored where available. We keep conflicts visible. For example, Mistral’s announcement and model documentation disagree on the active parameter count; we do not invent an explanation.
Cost calculation
Monthly cost = requests × (input tokens × input USD per million + output tokens × output USD per million) ÷ 1,000,000. We keep intermediate precision and round only for display.
Calculations assume the same input and output token counts across providers. They exclude cache pricing, batch discounts, taxes, tool fees and other billing tiers. DeepSeek peak and off-peak rates are separate records. The exported files include sources, dates and assumptions.
Weight memory calculation
Raw weight bytes = total parameters × bits per parameter ÷ 8. GB divides bytes by 10⁹; GiB divides by 2³⁰. This is a theoretical storage lower bound. It excludes runtime buffers, KV cache, quantization metadata and communication overhead.
We do not infer exact KV cache memory without verified architecture and serving details. We do not treat total GPU capacity as proof of a working deployment.
Benchmark rules
Different datasets, versions, harnesses and reasoning budgets are not merged into one ranking. A vendor’s attribution to an evaluator remains vendor-reported until we check that evaluator’s original record.
There are no site-run benchmarks in the first edition. Future tests must publish the model version, prompts, tools, settings, samples, failures and limits. Remote API tests and local weight tests will have separate labels.
Updates and corrections
Source records keep the check dates; those dates are not displayed throughout the pages. Price, license, artifact or framework changes require a source recheck and an updated deployment.
We do not promise a particular search ranking or deployment result. The job of this guide is to make the available evidence easier to inspect. Return to the overview.