It's 01:40 and the invoice-extraction worker is fine. The cost line in the job log is the problem: a few hundred PDFs, each one with a scanned delivery note stapled to it, and every call goes to the one model somebody picked in a sprint planning eighteen months ago. Nobody revisited the choice, because the model name lives in six places, two of them inside a prompt builder that also formats dates. That's the codebase I keep thinking about this week as Mistral ships Large 4. My position: the model itself is a respectable piece of work, but the release is most useful to us as an argument for building PHP apps where swapping the model is a one-line change in config and not a refactor.

Look at the numbers first, because they argue against picking Large 4 on general grounds. On the Artificial Analysis Intelligence Index it lands at 38. Xiaomi's MiMo-V2.6-Pro sits at 46, GLM-5.3 Max at 45, and Anthropic's Sonnet 5.5 and Opus 5.5 reach 56 and 58 at max effort. The sharpest comparison is GPT-6 Luna Max, which also scores 38 and costs roughly $0.07 per index task against about $1.13 for Large 4, using around 140 million output tokens where Mistral's model burns close to 200 million. Same score, about sixteen times the standardized bill. If your only question is which endpoint gives the most reasoning per euro, the answer is somewhere else, and I won't pretend otherwise.

That's the strongest objection, and it deserves full weight. Most of us building SaaS on Laravel or Symfony don't have a sovereignty clause in our contracts. We have a budget, a latency target and a product owner who wants the summary feature to stop hallucinating customer names. For that team, wiring in a 38-point model at the higher price is a bad trade, and loyalty to a European lab won't pay the invoice. Fair caveat on the caveat: the index uses Mistral's list rates of $1.36 per million input and $4.18 per million output tokens, while the public preview currently runs at half that, $0.68 and $2.09. Cheaper, still far from Luna territory.

Here's where I land differently. The aggregate score is an average over ten evaluations, and your application is not an average. Large 4 hits about 59.9% on AutomationBench, the multi-step email and spreadsheet style workflows, which puts it within a few points of GLM-5.3 Max at 62% and roughly level with MiMo at 59%. On Artificial Analysis' GDP.pdf document evaluation it reaches 19% where GLM-5.3 Max manages 11%. It accepts images natively through a 1.6 billion parameter vision encoder, while the GLM-5.3 that Mistral hosts is text-only. My delivery-note job cares about exactly that slice and nothing about competition math. A model can lose the leaderboard and still be the right call for one queue.

And notice what Mistral itself is doing. Its own platform already serves Z.ai's GLM-5.3, unmodified, next to the house flagship. The vendor is telling you, in product form, that the right model depends on the job and on the month. If the company that trained a 1.05 trillion parameter model, 49 billion active per step, on about 4,000 Grace Blackwell GPUs is comfortable routing customers to a competitor's weights, then our apps can be comfortable treating the model name as an environment variable. That means an interface for the call, a per-task mapping in config, and an eval fixture of twenty or thirty real inputs you rerun whenever a new release lands. Boring PHP. Exactly the kind we're good at.

Then there's the client who does have the clause. Public sector, health, an industrial supplier who won't let a diagnostic log leave the EU. For them, the open weights Mistral says will follow later this month are the actual news, because a model you can run on infrastructure you control changes the procurement conversation more than any index score. The jump from Medium 3.5, which scored 14 on the same index, to 38 is what makes that option usable for agent workloads at all. A year ago I'd have told such a client to accept a much weaker model or accept a foreign API. Now the gap is a cost and quality tradeoff you can measure on your own fixtures.

So I'm not telling you to switch to Large 4, and I'm not telling you to ignore it. I'm telling you the release exposes how many of our integrations can't even ask the question cheaply. If trying a new model means touching a prompt builder, a DTO and three service classes, you'll never benchmark it, and you'll keep paying for that sprint-planning decision forever. Make the model a value you can change, keep a small honest test set, and let each job pick its winner. Some months that will be Mistral, plenty of months it won't.

Where do you draw the line in your own stack: do you route per task with a mapping table and accept the extra prompt maintenance, or do you standardize on one provider and eat the cost difference for the sake of a simpler codebase?