Skip to content
le chonk
Menu

How to run Le Chonk locally

The weights are still pending a verified release. This is a preparation guide: what is available now, what you need to check, and how to verify a local run when an artifact is released.

On this page

Can you run the weights today?

Not from a release we have verified. The weights are pending a verified release; this guide has no confirmed weight download, final license or tested local inference command.

A public API preview is available. That does not mean the preview’s weights are downloadable or that a framework can load them. We will add a complete local tutorial after verifying the artifact and running the steps on named hardware.

There are no pretend terminal screenshots or commands marked “tested” on this page. A setup instruction is useful only if it matches an actual model release and supported environment.

Calling the API from your computer

You can write a program on your own computer that sends a request to Mistral’s servers. The program is local; the model inference is remote. You do not need enough local VRAM to hold the model for this use.

The model documentation lists mistral-large-4. Use the provider’s current API reference to confirm the request format, account access and available features. This guide has not executed an authenticated preview request, so it does not label an API example as tested.

Keep the key on your machine

Set your API key in your local environment or your app’s server-side secret store. This website never asks for a key and does not make model requests. Do not put a secret in frontend code, a public repository or a screenshot.

Measure one request before a batch

Check the returned model identifier, token usage and actual bill. Record whether reasoning tokens or other features affect the usage you see. The cost calculator is a planning tool, not a replacement for your provider’s usage records.

Loading weights on your own hardware

A complete local setup needs more than a model name. Confirm these items from the actual release before installing a stack or renting a server.

Local deployment readiness
Required itemCurrent statusWhat to confirm
Official weight artifactPending verified releaseDownload URL, repository owner, revision and file checksums
LicenseNot verifiedTerms for use, distribution and modified or quantized weights
Inference frameworkNot verified for this releaseExplicit architecture support in a specific version
Quantization formatNot verifiedAvailable artifact, supported kernels and quality checks
Recommended hardwareNo site-run testGPU models, per-device VRAM, host RAM and interconnect
Working launch commandNo site-run testExact environment, parameters and health check

Do not substitute a command for a similarly named older model. Framework support for one Mistral architecture does not establish support for a new release.

Prepare a memory budget

Use total parameters, not active parameters, for the initial weight calculation. At 1.05T parameters, raw 16-bit weights are about 2,100 GB; 8-bit is about 1,050 GB; 4-bit is about 525 GB. These are lower bounds with no runtime overhead.

Then budget separately for KV cache, framework buffers, quantization metadata and communication. Context length and concurrent requests affect memory use. Check host RAM, free disk space and transfer time as well as GPU capacity.

CPU or expert offloading may reduce GPU residency, but it can add transfer delays. A multi-GPU total does not prove that each device has enough free memory for its assigned tensors and runtime work.

Use the memory estimator and read its limits →

What a reproducible local run needs

Once weights and framework support are confirmed, a useful tutorial should let another person repeat the run. The checklist below is the planned verification sequence, not a set of tested commands.

  1. Record the environment. Operating system, GPU model and count, driver, runtime and framework version.
  2. Pin the artifact. Record the repository revision, file checksums, license and quantization.
  3. Install the supported versions. Save the dependency list and installation output.
  4. Start with the smallest useful configuration. Use a short context and one request before increasing the load.
  5. Check server health. Confirm the process is ready and that a model load did not silently fail.
  6. Send a first request. Save the prompt, settings, output and token usage.
  7. Measure the workload. Record cold-start time, first-token latency, output speed and peak memory.
  8. Repeat with realistic inputs. Increase context and concurrency gradually; include failures.

A successful response is a starting point. A usable deployment also needs enough speed and stability for the intended application.

If the setup fails, keep the useful details

Out of memory

Record which device ran out, the context and batch sizes, and whether failure happened during loading or generation. A load failure and a KV cache failure need different fixes.

Unsupported model or format

Check the exact artifact type and framework version. Do not assume an older release’s loader supports a new architecture. Keep the original error message and a minimal reproduction.

Slow generation

Separate cold-start time from steady-state speed. Check whether the runtime is offloading weights, waiting for transfers or sharing devices with another process. Report the configuration alongside throughput.

We will add specific fixes after reproducing real failures. Until then, this section gives a recordkeeping method rather than an untested troubleshooting recipe.

Sources & next steps