Home / Insights / Open-Weight vs Paid AI Models: Who Wins in 2026?
Insights

Open-Weight vs Paid AI Models: Who Wins in 2026?

Summarize with AI Prompt copied — paste it into the chat

Open-weight AI models are not beating the paid frontier on raw capability — they trail it by about four months — but that gap has stopped mattering for most business work. Epoch AI puts the lag at roughly four months, or eight points on its capability index, and it has not widened materially in two years. When a free model is one release behind and costs a fiftieth as much to run, "behind" stops being the deciding number.

Are open-weight models actually as good as paid models now?

For most tasks, yes. For the hardest tasks, not quite. The honest version has two numbers in it. Epoch AI's Capabilities Index put the best open-weight models about four months behind the closed frontier as of early 2026 — an eight-point ECI gap, comparable to the distance between two consecutive releases from a single lab. OpenRouter's June 2026 review found the same shape from the demand side: a consistent three-to-six-month gap sustained for over eighteen months, with no sign the frontier labs are accelerating away.

The specific models matter more than the average. On the Artificial Analysis Intelligence Index, GLM 5.2 led open weights at 51 in mid-June and topped the open field on Artificial Analysis's real-world agentic benchmark, landing effectively level with GPT-5.5 there. DeepSeek V4 Pro reached 80.6% on SWE-bench Verified, the top open-weight score, with the cheaper V4 Flash within about 1.6 points of it. In mid-July Moonshot AI launched Kimi K3 at 2.8 trillion parameters with roughly 104 billion active per token, shipping the full weights late in the month — the largest open-weight model publicly available.

So the frontier still leads. What changed is that second place became genuinely good, and second place is free to download.

What is the difference between open-source and open-weight AI?

Open weights means you can download and run the finished model. Open source means you also get the recipe — training code and enough detail about the data to rebuild and audit the model from scratch. Almost every model marketed as "open source" today is only open-weight. Meta, Google and most Chinese labs use the open-source label; the Open Source Initiative's Open Source AI Definition does not agree, because OSAID 1.0 requires training code under an OSI-approved licence, the parameters under free-use terms, and enough data information for a skilled person to substantially recreate the system.

This is not pedantry, because the licences differ in ways that reach your contracts. DeepSeek V4 Flash and GLM 5.2 ship under MIT, which is genuinely permissive. MiniMax M3's weights are out under a community licence requiring attribution, with larger commercial products needing prior written authorisation. Kimi K3 shipped under its own custom licence rather than MIT. NVIDIA's Nemotron 3 Ultra uses OpenMDW. "Open" is a spectrum, and the only version of the question that matters commercially is: what does this specific licence permit my company to do? Read it before it reaches a customer contract.

Why does this feel like the personal computer moment?

Because the personal computer did not win by beating the mainframe. Through the early 1980s a mainframe was faster, more reliable and better supported than anything on a desk. The PC won anyway, on three things the mainframe could not offer at any price: you owned it, it sat where the work happened, and it was good enough for the work most people actually did. Capability was never the deciding variable — control and proximity were.

The Netherlands has an unusually clear memory of exactly how that transition happened, because it was partly engineered by tax policy. From 1 January 1997 the PC-privéregeling let employers provide or reimburse a home computer free of tax and social premiums, up to 5,000 guilders of equipment. At its height, roughly one in six PCs reaching Dutch consumers was bought through an employer. Dutch manufacturers like Tulip Computers in Den Bosch — for a time one of Europe's largest — grew substantially on the back of those corporate PC schemes, and Philips put MSX machines in a generation of Dutch and Flemish homes. The scheme was abolished in August 2004, having done its work: it moved computing from something you requested from a central department to something that sat on your own desk.

The open-weight story has the same shape. A closed API model is timesharing: metered, remote, excellent, and someone else's. An open-weight model you host is the machine under the desk — a little behind on benchmarks, entirely yours, running where your data already lives. The analogy earns its place because the failure mode repeats too: plenty of firms took the PC-privé subsidy and never made those machines productive, because owning the hardware was never the hard part. Owning the process was, and that has not changed.

Where the analogy breaks down is worth stating. A PC was a one-off purchase that depreciated quietly; an open-weight deployment is a running commitment with GPUs, evaluation and updates attached. And the mainframe never got cheaper every quarter, while closed-model pricing keeps falling. Treat the analogy as a guide to the shape of the shift, not a prediction of who folds.

What does it actually cost to run an open-weight model?

Pull quote: The personal computer never beat the mainframe on power. It won because you owned it, and it sat where the work was. — Crux Digits

Two routes, two very different cost structures. Hosted, DeepSeek V4 Flash has been available under $0.10 per million input tokens and under $0.30 per million output — roughly two orders of magnitude below frontier output pricing. GLM 5.2 ran under $0.50 in and a few dollars out per million, still meaningfully under closed frontier coding models. The caveat that catches people: cheap tokens are not cheap runs, because these models think a lot, and a verbose reasoning model at a low token price can outspend a terse expensive one.

Self-hosted, the rule of thumb is about 0.5 GB of memory per billion parameters at 4-bit quantisation. That puts genuinely useful models on hardware you may already own: an Apple Silicon Mac with 32 GB of unified memory handles models in the 30B class at Q4, and Ollama, LM Studio and llama.cpp all run them without specialist skills. Vendors commonly claim self-hosting cuts total cost of ownership by up to 60% over three years versus API fees — treat that as a vendor figure and model your own, because it holds only at sustained high volume where the GPU stays busy.

When should a Dutch SME still pay for a closed model?

More often than the enthusiasm suggests. Pay when the task sits at the top of the difficulty range — hard multi-step reasoning, long-horizon agentic work — where a four-month gap is exactly the gap that matters. Pay when you need native image or video input and do not want to assemble it. Pay when nobody on the team wants to own an inference stack, which is the honest situation at most companies under fifty staff. And pay when the model is a small share of the bill anyway; below roughly a few hundred euros a month, self-hosting is a hobby with a business card.

The strategic answer is not to pick a side but to stay portable, which is the argument we made in AI Strategy Shouldn't Chase the Best Model and why the MCP tooling layer matters more than the model choice. Our comparisons of GPT-5.6 against Fable 5 and the Chinese open-weight models and of Claude Opus 5 both landed on the same conclusion: the model is the most replaceable component in your stack, so build so that replacing it is cheap.

Does the EU AI Act treat open-source models differently?

Yes, but the exemption is narrower than most summaries imply. To qualify, a general-purpose model must be released under a free and open licence permitting access, use, modification and distribution, with parameters publicly available, and it must not be monetised. That exemption then covers only two obligations: drawing up technical documentation, and supplying certain information to downstream providers. The training-data summary and copyright-policy requirements still apply, and the exemption falls away entirely for models designated as carrying systemic risk.

For an SME that deploys rather than trains models, the practical read is simpler: these obligations mostly land on the model provider, not on you. Your duties attach to what you build and how you disclose it. Choosing open weights does not exempt you from the AI Act — it mainly changes who upstream is carrying the documentation burden.

Why sovereignty makes this a Dutch and Belgian question, not just a cost question

In the Benelux the open-weight argument is landing less on price than on control, and 2026 has made that concrete. In April 2026 seven Dutch IT companies presented the Open Cloud Alliantie, pooling existing Dutch datacentres into a sovereign alternative to AWS, Azure and Google Cloud for hosting critical systems and sensitive government data under Dutch and European law. At EU level, a €180 million sovereign cloud tender was awarded to four European providers, Belgian operator Proximus among them. SURF has been piloting Nextcloud-based alternatives in Dutch education and research.

Open weights are the model-layer version of that same instinct. A model whose parameters you hold can run inside a Dutch or Belgian datacentre, under the AVG, with no transfer question to answer and no vendor able to deprecate it out from under you. That last point stopped being hypothetical in June 2026, when a US export-control directive forced Anthropic to take Fable 5 and Mythos 5 offline for every user worldwide. The order was lifted on 30 June and the models returned after nineteen days, so it ended well — but it arrived with no warning and no appeal. For a Dutch or Flemish firm with a process running on a single closed model, that is the risk that should concentrate the mind. Continuity, not cost, is the strongest argument for keeping an open-weight option warm.

What happens next?

Three things look reasonably safe to expect. The capability gap stays roughly where it is — four months has been stable enough, through several supposed step-changes, that betting on it widening has been a losing position for two years. Inference becomes the control point rather than the model: when weights are free, the margin and the lock-in move to whoever runs them well, which is why NVIDIA ships strong open models and why hosting providers compete hard on price. And the licence question gets sharper, not softer, as OSI works toward an update to OSAID and more labs ship near-frontier weights under bespoke terms.

The prediction we would actually stand behind is duller than the PC analogy suggests: most Dutch SMEs will end up running both, routing cheap high-volume work to an open-weight model and reserving the frontier for the hard tail, without ever making it a philosophical decision. That is what the PC era eventually looked like too — the mainframe did not vanish, it just stopped being where the interesting work happened. If you are weighing that split for your own stack, the decision is an evaluation exercise on your own tasks, not a benchmark-table exercise.

Model names, scores and prices in this piece reflect published figures as of 31 July 2026 and move quickly — verify against current sources before making a procurement decision.

Frequently asked questions

Which open-weight model should a Dutch or Belgian SME start with?

Start with whichever your hosting provider already serves well, then test on your own tasks rather than on benchmark tables. As a rough map: GLM 5.2 for planning and long-horizon coding, DeepSeek V4 Flash when cost dominates, MiniMax M3 if you need image or video input, and Nemotron 3 Ultra if a US-built model on the NVIDIA stack matters to procurement.

Is Llama open source?

Not by the Open Source Initiative's definition. Llama, DeepSeek, Qwen and Gemma are open-weight: you can download and run the parameters, but you do not get the training code and data information that OSAID 1.0 requires. Meta and Google use the open-source label anyway, so check the specific licence rather than the marketing term.

Can I run an open-weight model on my own hardware?

Yes, for small and mid-sized models. A rough guide is 0.5 GB of memory per billion parameters at 4-bit quantisation, so a 32 GB Apple Silicon Mac comfortably runs models in the 30B class. Ollama, LM Studio and llama.cpp make setup straightforward. Frontier-scale open models like Kimi K3 remain data-centre workloads.

Is it cheaper to self-host an open-weight model than to pay for an API?

Only at sustained volume. Hosted open-weight models are already extremely cheap — DeepSeek V4 Flash has run under $0.10 per million input tokens. Self-hosting wins when your GPUs stay busy and when data residency has its own value; below a few hundred euros of monthly API spend it rarely pays for the operational effort.

Does using an open-source model exempt us from the EU AI Act?

No. The open-source exemption applies to providers of general-purpose models, covers only two obligations, and disappears for models with systemic risk — training-data summaries and copyright policies still apply. If you deploy rather than train models, your obligations attach to what you build and how you disclose it, regardless of which model you chose.
Our AI services Hire an AI consultant AI automation AI agents AI implementation Pricing

Want any of this applied to your business?

We turn these concepts into working tools — grounded, safe and measurable. Start with a free consultation.

Book a free consultation →