· 9 min read
Open Weights Will Win the Model Layer, Even if Closed Models Keep the Frontier
- artificial intelligence
- large language models
- open weight
- enterprise software
On April 5, 2025, Meta made the weights for its Llama 4 Scout and Llama 4 Maverick models available for download, allowing other companies to run them outside Meta’s own service. That was a product shipment, not a research promise, and it captured the economic question now facing large language models: once a capable model can leave its maker’s servers, what exactly remains scarce?
By July 2031, open-weight models are 70 percent likely to provide the default foundation-model supply for applications that do not require the very best model available that month. I am not claiming that every leading model will be open-weight, that model builders will stop earning money, or that a downloadable model is automatically safe. The prediction is narrower: the model itself will become an interchangeable input in most deployed systems, and open weights will be the format best suited to distributing that input.
Definitions matter before the forecast does
Open-weight means that a developer can obtain the trained parameter files needed to run inference on infrastructure the developer controls. The license may still impose conditions. Meta’s Llama licenses, for example, are not equivalent to an unrestricted open-source license. For this essay, the decisive feature is possession of runnable weights, not access to training data, code, or an API.
Closed-weight means the customer can use a model only through infrastructure operated by its maker or an authorized service provider. The customer buys outputs, rate limits, and a contract, but cannot independently host the model. This definition holds whether the interface is a consumer chatbot, a cloud endpoint, or software embedded in another product.
Commodity also needs a tighter meaning than “cheap.” A model is commoditized for a task when several suppliers can meet the task’s quality threshold, the buyer can switch suppliers at modest cost, and performance differences no longer justify a large price premium. Frontier research can remain rare while production inference becomes a commodity. Those are separate markets.
The force behind the shift is a collapsing cost curve plus a change in distribution
The central driver is not a cultural preference for openness. It is the separation of a costly creation process from a cheap-to-copy artifact. Training a frontier model still requires unusual concentrations of chips, electricity, data engineering, researchers, and time. A finished weight file, by contrast, can be replicated to another host at negligible marginal cost. The first activity will remain concentrated longer than the second. But most buyers purchase the second activity: tokens generated for a particular workflow.
Inference economics make the separation sharper. Stanford’s 2025 AI Index reported that the price of a model meeting the GPT-3.5-level threshold on MMLU fell from $20 per million tokens in November 2022 to $0.07 per million tokens in October 2024. That is a 280-fold decline over roughly 23 months for a fixed quality bar. The figure is a lagging indicator because it records prices after providers have already built and deployed better systems. Its importance is that any model which clears a stable business threshold faces rapid price pressure thereafter.
Open weights turn that price pressure into a market structure. A company can choose its own cloud, use a specialist inference host, run a model inside its own network, quantize it for cheaper hardware, fine-tune it on internal examples, and keep working if a model maker changes an API price or policy. Each option weakens the original model provider’s control over the buyer. The customer does not need every option. The credible ability to leave changes negotiations.
The causal chain is straightforward. First, a closed provider demonstrates a new capability and attracts early demand. Second, competitors reproduce enough of that capability through improved training recipes, synthetic data, distillation, architecture changes, or more focused post-training. Third, an open-weight release gives infrastructure providers and application companies a common target for optimization. Fourth, those optimizations reduce serving cost and improve integration. Fifth, buyers whose tasks clear a quality threshold move from paying for a named model to buying the least expensive reliable system that meets their requirements. The premium concentrates at the temporary frontier, while the much larger volume of routine work becomes contestable.
Distribution accelerates the process. An API is a product channel controlled by one supplier. Weights are a component that can be distributed by cloud companies, hardware vendors, systems integrators, device makers, and internal platform teams. Once the component is broadly available, every participant has an incentive to make it easier and cheaper to use. Cloud providers want compute consumption. chip vendors want optimized workloads. Enterprises want control over data and uptime. Application builders want lower unit costs. None needs the original model maker to retain exclusive control of the model.
This is why Meta’s strategy is economically coherent even if it earns less directly from model access than a closed provider would. A company with a large consumer distribution business can benefit when the cost of generic intelligence falls across the industry. It can spend on the frontier, release weights, and shift value toward the complements it already owns: user attention, distribution, advertising inventory, devices, and developer adoption. The model maker that sells only model access has a different incentive. It needs the model layer itself to remain scarce.
A leading indicator is the widening availability of capable downloadable models in multiple sizes and licenses. Mistral released Mistral Small 3.1 under Apache 2.0 on March 17, 2025, and said it could run on a single RTX 4090 or a Mac with 32 GB of memory. Meta reported more than one billion Llama downloads on April 29, 2025. Vendor-reported downloads are imperfect measures of production use, but they do show that weights have become a distribution format with independent demand. The better lagging indicators will be enterprise procurement data, hosting revenue, and token volumes, which arrive later and are usually disclosed selectively.
The historical analogy is Linux, with an important limit
The useful analogy is not that open-weight models will eliminate proprietary software. Linux did not erase proprietary operating systems or make operating-system vendors irrelevant. It did, however, become a widely deployed substrate on which cloud providers, hardware companies, service firms, and application vendors could build. Much of the economic surplus accrued around the kernel rather than to whoever distributed the kernel.
The same split can happen with language models. Adoption of an open component can rise while profit moves to the companies selling chips, managed inference, private deployment, proprietary data connections, workflow software, security reviews, and human support. That is the case where adoption goes one way and profit another. Open weights can win distribution without creating a durable standalone profit pool for the organizations that publish them.
The analogy breaks down at the frontier. An operating system is a relatively stable interface. A frontier model is an evolving capability whose reliability, tool use, multimodal inputs, safety behavior, and reasoning can all change materially between releases. A model provider can still command a premium when it delivers a capability others cannot reproduce, especially if that capability improves revenue or reduces headcount in a high-value workflow. The open-weight case is strongest below that moving frontier, where buyers value control and cost more than a temporary quality edge.
The strongest case for closed weights
Dario Amodei, Anthropic’s chief executive, makes the best version of the opposing argument. His position is not simply that closed models are better products. It is that increasingly capable models create security and misuse risks that require strong protections around the weights and around deployment. Anthropic’s Responsible Scaling Policy treats the ability to secure model weights as part of the response once a system reaches sufficiently dangerous capability levels.
That argument has real force. An API provider can update a safeguard centrally, monitor suspicious use, revoke access, apply a patch after a failure, and invest in a costly evaluation program without depending on every downstream deployer to do the same. Closed deployment also gives customers a single accountable counterparty. For a bank, hospital, defense contractor, or company using a model in a consequential decision process, that accountability can be worth more than the savings from self-hosting.
Closed providers also have a commercial advantage when the product is more than a base model. A system that combines a model with proprietary tools, retrieval, long-running memory, enterprise connectors, evaluation data, and a managed agent runtime can be difficult to reproduce merely by downloading weights. The buyer may be paying for operating discipline and product integration rather than for raw text generation. This is where the claim that “models are commodities” becomes too broad.
The closed-weight view gets another point right: an enduring frontier gap would change the forecast. If the best closed systems keep a performance lead large enough to alter business outcomes for more than a year, then open-weight alternatives will be substitutes only in lower-stakes work. The prediction here assumes that replication, distillation, and specialization continue to shrink that gap for mainstream tasks. It does not assume that frontier leadership becomes permanent or irrelevant.
Where durable profit is likely to settle
The likely outcome is a barbell. A small group of companies will spend heavily to produce the next frontier capability and charge for early access. A much larger market will run open-weight or open-weight-derived models tuned for defined jobs: customer support, document extraction, internal search, coding assistance, translation, classification, and structured workflow steps. The middle, where a general-purpose API charges a large premium for capability that several other models can match, is the vulnerable position.
That does not make the model unimportant. Electricity is not unimportant because it is standardized, and databases are not unimportant because several vendors sell them. It means competition shifts from ownership of a model artifact to the surrounding system: data rights, distribution, latency, reliability, compliance, product design, and the authority to act inside a customer’s workflow. A company that mistakes a fast-improving model for a permanent tollbooth will discover that its customers were buying a result.
Two tests for the forecast
First prediction: by July 23, 2028, at least one downloadable model under a permissive commercial license will score within 10 percent of the leading closed model on a publicly documented, independently run evaluation of tool-using coding work, while costing less than half as much per completed task on comparable hardware. If no open-weight model reaches that band by that date, the convergence claim is weaker than this essay assumes.
Second prediction: by July 23, 2031, most large enterprises that run generative AI for high-volume internal tasks will operate at least one open-weight model through their own infrastructure or a managed host that permits workload portability. I would admit the broader forecast is wrong if closed providers preserve a measured performance advantage of more than 25 percent on mainstream enterprise tasks for three consecutive years, or if regulated buyers overwhelmingly reject portable weights even after comparable models are available. The concrete price to watch remains the cost per completed task, not the number of model announcements.