· 8 min read
DeepSeek’s Distillation Test for the Frontier-Model Tollbooth
By S. Huang
- artificial intelligence
- deepseek
- anthropic
- openai
- model distillation

On February 23, 2026, Anthropic said that DeepSeek, Moonshot and MiniMax had generated more than 16 million exchanges with Claude through roughly 24,000 fraudulent accounts, alleging that the activity was used to extract capabilities for competing models. Anthropic’s claim remains an allegation by an interested party, but its scale identifies the commercial issue clearly: a frontier model can become a data-generating teacher for a cheaper competitor.
The prediction is that, by December 31, 2028, distillation will cut the share of Anthropic’s and OpenAI’s revenue that can be defended by general-purpose model intelligence alone, with a 70 percent probability. This is not a claim that either company will lose its technical lead, that every model can be copied from API outputs, or that DeepSeek has proved Anthropic’s allegation. It is a claim about bargaining power: once a customer’s recurrent workload can be taught to a smaller model, the frontier provider’s per-token price and margin face a ceiling.
What distillation means here
Distillation, in this essay, means training a student model on a teacher model’s input-output behavior so the student can reproduce useful performance on a defined distribution of tasks. The definition includes visible answers and intermediate reasoning traces when those are available. It does not mean stealing weights, replicating every capability, or independently recreating the teacher’s training process. That distinction matters because output-based distillation can be legal and ordinary within one company, licensed and explicit between parties, or prohibited when a competitor uses API output against the provider’s terms.
DeepSeek demonstrated the first, legitimate version in public. Its January 2025 R1 release included distilled Qwen models at 1.5 billion, 7 billion, 14 billion and 32 billion parameters, plus Llama-based versions, and said the Qwen variants were fine-tuned on 800,000 samples curated with DeepSeek-R1. The repository used an MIT license for the R1 weights and expressly allowed commercial use, modification, derivative works and distillation. That was a shipped example of a large reasoning model being converted into smaller deployable models, not merely a research proposition.
The mechanism: outputs turn a service into training input
The force driving this trend is a change in the cost structure of capability transfer. Training a frontier model requires vast compute, data work, experimentation and evaluation. Distilling a task-specific student starts after those costs have already been paid by somebody else. The student builder buys or obtains a set of high-quality teacher responses, selects the prompts that matter to a target product, fine-tunes an existing base model, and serves that narrower model at a lower inference cost. Each step removes a portion of the original laboratory expense from the follower’s balance sheet.
That chain is especially damaging to the API business model because the most profitable use cases are often repetitive. A software company may need a frontier model to solve a hard class of coding, support, document-processing or agentic tasks during development. Once it has identified a stable prompt distribution and enough accepted outputs, it can train a student for that distribution. The frontier API changes from a permanent production dependency into a temporary teacher and evaluator. OpenAI itself described precisely this economic logic in its October 1, 2024 distillation product announcement: outputs from a more capable model can fine-tune a smaller model that matches performance on specific tasks at lower cost.
DeepSeek adds a second pressure point. An openly distributable teacher does not merely compete for inference requests. It lowers the cost of producing the next generation of students for everyone able to fine-tune a base model. Hugging Face’s Open-R1 project published a recipe to reproduce the reasoning capabilities of DeepSeek-R1-Distill-Qwen-7B, reporting closely comparable benchmark results for its own 7-billion-parameter student. That is a leading indicator. It does not measure revenue lost by Anthropic or OpenAI, but it shows that the practical know-how, recipes and evaluation targets can spread beyond the original model maker.
The lagging indicator will be different: declining realized revenue per unit of routine work, after controlling for demand growth, at the frontier labs. Public model launches and benchmark scores are leading indicators because they show that substitution could occur. A customer moving a stable workload from Claude or OpenAI to a self-hosted or lower-priced student is the event that changes economics. Neither company’s headline valuation nor total usage alone answers that question, because total usage can rise while the highest-margin calls migrate to cheaper models.
Why contractual barriers are weaker than technical moats
Both companies understand the threat. OpenAI’s November 2023 business terms prohibited model extraction or stealing attacks and prohibited using output to develop AI models that compete with its products and services. Anthropic, in February 2026, said it had detected and disrupted what it called industrial-scale distillation campaigns and framed the activity as a terms violation. Those rules can deter identified customers, support account enforcement and create a factual record for litigation. They do not change the underlying incentive: if a teacher is materially better and a student can capture enough of its behavior, a large price gap pays for evasive acquisition efforts.
That is why abuse detection matters more than legal wording. A provider needs to identify coordinated accounts, query patterns designed to map capability boundaries, anomalous volumes, answer harvesting and attempts to collect reasoning traces. It then needs to limit those patterns without making normal enterprise automation unusable. The defense is inherently costly because a useful API is meant to answer many prompts at scale. The more broadly a provider sells access, the more valuable its outputs become as potential training material.
The IBM PC analogy, and where it fails
The closest historical analogy is the IBM PC’s open architecture, not because language models are PCs, but because the economic pattern is similar. IBM says its 1981 PC used off-the-shelf components and published technical reference material. That openness helped establish a standard and brought software and peripheral makers into the market; it also enabled compatible competitors. IBM later said that clones had eroded its PC dominance by 1986. The lesson is that adoption of a technical standard can expand the total market while weakening the originator’s claim on the finished product’s profit pool.
The analogy breaks in important ways. A PC’s interfaces were published specifications, while a closed model provider can change weights, tools, policies, rate limits and output behavior without publishing internals. Model outputs are probabilistic, and a student trained from them will usually inherit blind spots. Frontier labs also own distribution through consumer products, enterprise contracts, developer tooling, safety work and proprietary feedback loops. IBM’s problem was compatible hardware around a relatively stable standard. Anthropic and OpenAI confront an opponent that can copy slices of behavior while the target keeps changing.
That difference creates the strongest opposing case. Anthropic’s position, expressed in its February 2026 disclosure and policy work, is that distillation attacks can be detected, accounts can be disrupted, advanced-chip controls can constrain scaling, and frontier capability remains costly to acquire independently. Its smartest version is not that output extraction is impossible. It is that a copied student will lag at the moving frontier, fail on rare but valuable tasks, and leave serious enterprises willing to pay for a provider that offers stronger models, reliability, indemnities, security controls and rapid upgrades. Anthropic gets a great deal right. Distillation is much better at compressing known, frequent work than at reproducing broad judgment, novel research ability or an evolving product stack.
The weakness in that defense is economic rather than technical. A student does not need to beat Claude or OpenAI’s best model in the abstract. It only needs to clear the customer’s quality threshold on enough volume to move spend. If a customer keeps the frontier API for exceptional cases but routes 80 percent of stable work to an internal or open student, the frontier lab can retain prestige and grow total tokens while losing the work that once subsidized its frontier training bill.
Two tests before the end of 2028
First prediction: by June 30, 2027, either Anthropic or OpenAI will offer an explicitly managed pathway for customers to create task-specific smaller models from frontier-model outputs while keeping the resulting data and deployment inside that provider’s commercial boundary. I would admit this prediction is wrong if neither company has announced such a product or contractual program by that date. OpenAI’s 2024 distillation workflow is an early version of this strategy, so the test is whether the major providers make it a central enterprise retention product rather than a specialist feature.
Second prediction: by December 31, 2028, at least one major frontier lab will disclose, in a product policy, security report or earnings-related statement, that automated extraction attempts have forced material limits on model-output access or reasoning-trace availability for some customer tier. I would admit this prediction is wrong if providers continue to offer unrestricted high-volume output access without reporting such a restriction, while independent student models fail to achieve commercially useful parity on recurring coding or document workflows. The concrete measurement is a published restriction, not another benchmark chart.