The smartest AI model outside the U.S. and China comes from France. Mistral Large 4, announced as the playful “Le Chonk” with a chubby cat as its mascot and boasting over 1 trillion parameters, will be available to download for free starting at the end of this month. Running it yourself does require massive compute given the constraints, but as an open-weight offering there will be plenty of providers to pick from. For now, it’s available for testing in preview. Although this LLM isn’t a competitor to the best models from OpenAI, Anthropic, and Google, this release puts Mistral back on the map in the AI race.
That relevance depends on several factors. For example, AI benchmarks, as enticing as they may be, have increasingly proven to be less of a reflection of actual performance. Nevertheless, the score of 38 on Artificial Analysis is a good indicator: this is just below the level of DeepSeek V4.1 Flash (39), a model that many users view as a cost-effective yet powerful option for AI workloads. Mistral Large 4 does seem more expensive per task. However, if you take Mistral’s own claims at face value, you’ll see that Mistral performs coding and security tasks significantly better than DeepSeek.
When comparing performance to the AI giants OpenAI and Anthropic, GPT-6 Luna emerges as the closest competitor. That is OpenAI’s smallest model; GPT-6 Terra, 6.1 Sol, 6 Astra, Claude Sonnet 5.5, and Opus 5.5 are significantly more powerful, with scores in the 50–60 range instead of 38, to use another “ballpark figure.”
The “vibes”
Anyone who, for whatever reason, is limited to open-weight models can freely choose between providers. Platforms like OpenRouter make it possible to shift tasks from one LLM to another with just a few clicks. That makes Mistral Large 4 just one of many tools in the developer’s or end-user’s arsenal. Sometimes Moonshot AI’s Kimi K3 from China is the most powerful option, despite the hefty cost due to its 2 trillion parameters, but often the specific use case will determine which LLM performs best.
Ultimately, it takes a few days to get a feel for the “vibes,” that is, the subjective experience of what a model can and cannot do well. For example, models like Meta Llama 4 and, more recently, Claude Opus 5 turned out to be far less capable and consistent than initially thought, while the series of releases from Anthropic and OpenAI over the past few months has continually led to a reassessment of both providers.
Two worlds
Mistral Large 4 may compete with open-weight models from China, but for most organizations, OpenAI and Anthropic are the most obvious options. Here, too, there are various formats and price points to choose from depending on the objectives.
Model routing has not yet fully crystallized. It’s clear that selecting a specific model with a specific “effort” level for each prompt disrupts the ideal workflow. The point of light AI usage as an alternative to traditional UI navigation is that it involves less friction, so ideally, every prompt would automatically include the correct model and effort selection. Instead, many users simply choose the best available model and stick with it. This works better for Anthropic and OpenAI than for the competition because their subscription options are still significantly cheaper than API usage. Those who push the limits of Claude Max or ChatGPT Pro to the max get much more value for their money than with open-weight models, regardless of the latter’s quality.
However, organizations that opt for Mistral also have a choice between large and small models, including options that run locally even on conventional workstations. The top and base levels of these models, however, are lower than those of the American AI giants. This creates two worlds that are difficult to combine, although Claude and GPT can also be accessed via the API, this isn’t particularly attractive within OpenRouter due to the steep prices.
Once at the top, now affordable
Organizations do need to ask themselves what they actually need from their AI tools. If their goals were already being met with, say, Claude Opus 4.5 at the end of last year, then Mistral Large 4 is significantly more intelligent, faster, and cheaper. For that reason, LLMs at a lower price point quickly justify their existence as replacements for earlier models. The problem organizations may encounter is that their expectations grow alongside the quality improvements of the LLMs in question. Opus 5.5 is, in fact, much more powerful than 4.5, 4.6, 4.7, 4.8, or 5 ever was, and many developers will have become familiar with each incremental improvement and grown accustomed to it. As a result, organizations are delegating more tasks to LLMs, trusting the outputs more often, being more ambitious with their workflows, and accepting the higher price. If that dynamic doesn’t hold or the improvements stagnate, companies like Mistral and DeepSeek will be able to beat Anthropic and OpenAI’s value for money. For now, that hasn’t happened.
Read also: Claude Opus 5.5 released: better and cheaper than Fable 5.1