Blog Post

Mistral Large 4 (Le Chonk): Preview, Pricing, and the Open-Weight Catch

Mistral's trillion-parameter Le Chonk is in API preview. Verified benchmarks, launch pricing, release-date differences, and what self-hosting would require.

Mistral Large 4 (Le Chonk): Preview, Pricing, and the Open-Weight Catch: Blog post featured image

Mistral Large 4, nicknamed Le Chonk, entered public API preview on October 6, 2026. Public weights are still forthcoming. That distinction matters if you are choosing a model for private infrastructure today. Mistral's release changelog confirms both the preview and the pending weights.

Our read: Le Chonk deserves a workload-specific evaluation, especially for security and document-heavy products. Its trillion-parameter size makes self-hosting a serious infrastructure decision. We would start with the API, measure completed tasks, and make the deployment decision when the checkpoint and its terms arrive.

Sources checked October 7, 2026. We have not run our own model evaluation; the numbers below are attributed to their publishers.

What is Mistral Large 4, and why do 49B and 52B both appear?

Large 4 is a general-purpose multimodal Mixture-of-Experts model. Its model documentation lists 1.05 trillion total parameters, a 1.6-billion-parameter vision encoder, and a 1M-token context window. The API model ID is mistral-large-4; supported features include structured outputs, function calling, document Q&A, batching, and agents.

The versioned model page says 49B active parameters, while the general Large page says 52B. Mistral's official Hugging Face listing explains the accounting: 49B active per token, or 52B including embeddings and output layers. These counts should not be presented as an uncertain range.

Sparse activation reduces the computation applied to each token. It does not make a trillion-parameter checkpoint occupy the memory of a 49B dense model.

Open weights are planned for month-end; the exact date differs

The English announcement promises weights by the end of October. The French announcement names October 26, while the official Hugging Face release listing displays October 31. October 27 should not be repeated as an official release date.

Plan around month-end, and recheck the actual checkpoint before scheduling a migration. The pages reviewed do not establish final public weight-license terms, so we cannot substantiate a specific custom license or assume Apache 2.0.

For a team with a fixed delivery date, API availability and downloadable weights are separate milestones. An open-weight roadmap gives you a possible deployment path; your production plan still needs the files, serving support, and usable terms.

Mistral Large 4 benchmarks show a stronger security case than an overall lead

Artificial Analysis's own launch evaluation gives Large 4 an Intelligence Index score of 38, equal to GPT-6 Luna at maximum reasoning effort and close to DeepSeek V4.1 Flash at maximum effort, which scores 39. Its Cyber Index score is 50, matching GLM-5.3-Flash and trailing MiMo-V2.6-Pro at 56.

Selected preview results, checked October 7, 2026
EvaluationResultSource
Intelligence Index38Artificial Analysis
Cyber Index50Artificial Analysis
CyberGym-E2E-AA82%Artificial Analysis
Cybench93%Mistral-reported
DeepSWE v1.161.7%Mistral cites Artificial Analysis
Terminal-Bench 428.3%Mistral cites Artificial Analysis

Sources: Artificial Analysis's independent results and Mistral's capability breakdown. Mistral attributes the two coding scores to Artificial Analysis; we verified that attribution in the announcement, rather than independently reproducing the runs.

The coding results are a useful brake on a broad superiority claim. DeepSWE and Terminal-Bench measure different workloads, so the percentages are not directly comparable. They are a reason to test repository changes and terminal work separately. A security win should not automatically make Le Chonk your default coding model.

Strong cybersecurity performance does not mean an unrestricted public API

Mistral describes reduced-moderation access for vetted cybersecurity partners and state authorities. It also reports strong refusal behavior on malicious cyber requests in its safety evaluation. The public preview should not be described as universally refusal-free.

For a defensive product, test the specific authorized tasks you intend to support: vulnerability triage, patch verification, and detection-rule drafting. Record both successful completions and refusals. A benchmark score does not tell you whether your account's endpoint will accept a particular investigation.

Mistral Large 4 pricing: budget beyond the launch discount

The official pricing page publishes these USD rates per million tokens:

Token typeStandard rateLaunch rate
Input$1.36$0.68
Cached input$0.14$0.07
Output$4.18$2.09

The changelog specifies 50% off for two weeks. Use standard rates for a recurring budget.

For an illustrative request with 100,000 uncached input tokens and 10,000 output tokens, our calculation is $0.0889 at launch rates and $0.1778 at standard rates. Ten identical calls cost $0.889 or $1.778, before tools, retries, taxes, or other charges. Actual agent calls will have different token counts.

Artificial Analysis estimates cost per Intelligence Index task at $1.13 using standard pricing, compared with $0.25 for GLM-5.3-Flash and $0.27 for DeepSeek V4.1 Flash at maximum effort. That is a measured benchmark comparison, not a quote for your application.

If you are evaluating sovereignty and price together, put an explicit value on control. Then compare cost per accepted result. Token rates alone can hide a model that needs longer answers or more attempts.

A 1M context specification is not a verified endpoint limit

Mistral's model card lists 1M context. The Artificial Analysis preview model page lists 524K. We could not establish the reason for that difference from those pages.

Verify the limit on the route your application uses. For document analysis, test whether the model finds known evidence near the beginning, middle, and end of your corpus, and whether its citations point to the right pages. Capacity and reliable recall are separate measurements.

European infrastructure changes the deployment question, not the memory math

Mistral says it trained Large 4 on 3,800 NVIDIA Grace Blackwell GPUs in its own European datacenters and serves the preview on that infrastructure. Official announcement.

Our weight-only estimate for 1.05 trillion parameters is about 2.1 TB at 16 bits, or 525 GB at 4 bits, using decimal units. That multiplication excludes quantization metadata, context caches, activations, and runtime overhead. It is not a hardware recommendation or evidence that a supported 4-bit checkpoint exists.

Self-hosting could give your team control over upgrades and serving policy, but the operating bill needs to include idle capacity, concurrency, and on-call work. This connects to our earlier vendor-dependency discussion: deployment control is useful when your team can actually operate the alternative.

What to test during the preview

Start in Mistral Studio with a held-out set of tasks from the product you already run. We would include ordinary requests, difficult exceptions, and examples your current model fails.

  • Score the finished artifact. Check patches against tests, extracted figures against documents, and citations against source pages.
  • Keep the comparison stable. Use the same tools, acceptance criteria, and spending limits for your baseline and Le Chonk.
  • Capture the full run. Record cost, completion time, retries, tool errors, and refusals, alongside task success.
  • Revisit self-hosting after release. Inspect the actual license, checkpoint formats, and serving guidance before reserving hardware.

Bring that evaluation to your next architecture review. The useful question is whether Le Chonk improves a workload enough to justify its API bill today and its operating cost tomorrow.

Explore More Articles

Discover other insightful articles and stories from our blog.