
AI Model Distillation: Why Does Garry Tan Want U.S. AI Labs to Copy Frontier Models?
Imagine building a smaller AI model by repeatedly asking a much larger model questions until you learn enough about how it responds to create your own system. That basic idea sits at the center of the AI model distillation debate now being pushed into the spotlight by Y Combinator CEO Garry Tan.
Tan says U.S. regulators should not necessarily stop American AI companies from using distillation techniques on frontier models. Instead, he wants smaller U.S. open-weight AI labs to be able to use similar techniques, arguing that this could create a stronger American ecosystem of open models rather than leaving that space increasingly dependent on Chinese AI companies.
His position is controversial because distillation itself is not inherently illegitimate. AI labs commonly use it as a training technique, but Anthropic has accused Chinese AI labs of conducting “illicit distillation attacks” involving concealed identities, fraud, and stolen credentials.
Tan’s proposal is different from that alleged behavior. He is arguing for legitimate, authorized access through the front door rather than unauthorized access through stolen credentials.
So why does this matter?
Because the argument is no longer simply about one AI training technique. It touches three much bigger questions: Who should control access to AI intelligence? How open should powerful AI models be? And could restrictions designed to protect frontier AI companies accidentally reduce competition?
What Is AI Model Distillation?
Question: What is AI model distillation?
AI model distillation is a technique in which developers use a more capable model to help train or develop another model, often by repeatedly prompting the original system and learning from its outputs.
The basic concept is easier to understand with a teacher-and-student analogy.
Imagine an experienced teacher solving thousands of difficult problems. A smaller student model observes the teacher’s answers and reasoning patterns and learns to produce useful responses of its own.
The student does not necessarily receive the teacher’s internal model parameters or source code. Instead, it can learn from the information the teacher makes available through interactions.
This can make distillation a powerful way to develop smaller or specialized AI systems.
Definition + Expansion: AI Model Distillation
AI model distillation is a model-training approach that transfers useful capabilities or behavioral information from a more capable “teacher” model to another “student” model.
Distillation can be legitimate and widely used in AI development. Problems arise when developers obtain access through unauthorized methods or deliberately circumvent restrictions imposed by the model provider.
That distinction is at the heart of the current debate.
A company could, for example, have legitimate access to an AI model through an API and use its outputs as part of a permitted training process. That is fundamentally different from obtaining someone’s credentials, disguising the identity of the user, or deliberately bypassing access controls.
Tan’s argument concerns the first kind of scenario.
What Exactly Is Garry Tan Proposing?
Question: What does Garry Tan want U.S. AI companies to do?
Garry Tan wants smaller American open-weight AI labs to have greater freedom to use distillation techniques on American frontier AI models, rather than allowing restrictive rules to prevent them from learning from models they can legitimately access.
Tan discussed the issue with CNBC and TechCrunch in September 2026.
When CNBC asked about Chinese AI labs using distillation, Tan said he would do nothing and suggested that there could instead be an “American distillation regime.”
He later clarified his position to TechCrunch.
His argument is not that American companies should steal credentials or secretly access models. Instead, he believes they should be allowed to use AI systems through legitimate access.
That is an important distinction because the debate can otherwise become confusing.
Tan is essentially asking whether a company that legally provides an AI service should be able to control every downstream use of information produced through that service.
His answer appears to be: not necessarily.
He believes excessive restrictions could make it harder for smaller AI companies to compete with the organizations operating the largest frontier models.
Why Does Distillation Matter in the U.S.-China AI Race?
Question: Why has AI model distillation become a geopolitical issue?
The debate has become geopolitical because Anthropic has accused Chinese AI labs of using illicit distillation techniques to extract knowledge from frontier models.
Anthropic released its second report on the issue in September 2026. The company alleged that Chinese labs were conducting “illicit distillation attacks” while hiding their identities and relying on fraud and stolen credentials.
Anthropic CEO Dario Amodei had also previously called for U.S. regulators to take action against distillation.
Tan takes the opposite position.
Instead of restricting the technique broadly, he wants American AI companies to be able to use similar capabilities legally.
This reflects two different approaches to AI competition.
One approach prioritizes protecting frontier AI companies from having their capabilities replicated by competitors.
The other prioritizes creating a competitive ecosystem in which smaller companies can learn from and build around powerful AI systems.
Neither position is trivial.
Frontier AI companies spend enormous amounts of money developing models and may view unrestricted distillation as a threat to their competitive advantage.
Open-weight AI companies, meanwhile, need ways to compete without having access to the same capital, compute resources, and research teams as the largest AI labs.
The policy question is therefore about finding the balance.
What Is the Difference Between Legitimate and Illicit Distillation?
Question: Is all AI model distillation illegal or unethical?
No. Distillation is a legitimate AI development technique, but the way a model is accessed and the terms governing that access matter.
A simple distinction looks like this:
| Type | How it works | Main issue |
| Legitimate distillation | A developer uses authorized model access and permitted outputs | Generally part of normal AI development |
| Unauthorized distillation | A developer violates a provider’s access restrictions | Can breach contractual or platform rules |
| Illicit distillation attack | A developer allegedly hides identity, uses fraud, or stolen credentials | Raises serious security and legal concerns |
| Open model training | Developers train or adapt models using available weights/data | Depends on the model’s license and data rights |
This is why Tan’s position should not be summarized simply as “he wants American AI companies to steal models.”
That would misrepresent the argument presented in the TechCrunch report.
His proposal is specifically about allowing U.S. labs to enter through legitimate channels.
At the same time, his argument raises a difficult question: how much control should a model provider have over what customers do with the information its model generates?
Why Does Garry Tan Think AI Intelligence Should Be More Open?
Question: Why does Tan believe access to AI capabilities should be less restrictive?
Tan argues that powerful AI systems were themselves developed using a vast amount of publicly accessible human knowledge, and he questions whether the intelligence produced from that knowledge should remain entirely locked behind restrictive terms of service.
He pointed to the history of AI training as part of his argument.
Proprietary AI companies have trained models using enormous quantities of human-created material. Some of that material has been copyrighted, and AI companies have faced significant legal disputes over how such material was used.
Tan therefore sees an asymmetry.
If frontier AI companies can learn from broad collections of human knowledge, should smaller AI companies be prevented from learning from the outputs of those frontier systems when they have legitimate access?
That is the philosophical foundation behind his position.
It is not simply about whether one company can reproduce another company’s technology.
It is about whether access to machine-generated intelligence should increasingly be treated as a public resource, commercial service, or tightly controlled intellectual property.
Those are very different policy frameworks.
Open-Weight AI Models vs. Frontier Closed Models
Question: Why does Tan care so much about open-weight AI models?
Tan believes open-weight models can provide people and companies with greater freedom and access, while frontier proprietary models can continue pushing the technological boundary.
He described the two sides as complementary.
Frontier labs are responsible for developing increasingly powerful systems. Open-weight labs can then provide alternative models that developers can inspect, adapt, and deploy with fewer restrictions, depending on their licenses.
The difference can be summarized this way:
| Feature | Frontier closed-weight models | Open-weight models |
| Model weights | Generally not publicly released | Weights may be publicly available |
| Control | Primarily held by model provider | More control can be available to developers |
| Customization | Often restricted to provider’s tools | Can be broader depending on license |
| Access | Usually through hosted services or APIs | Can potentially be deployed independently |
| Business advantage | Strong provider control | Broader developer ecosystem |
| Main concern | Concentration of AI power | Misuse and safety challenges |
Open-weight does not automatically mean “completely open.”
Licenses can impose conditions, and developers may still face restrictions around commercial use, safety, redistribution, or other activities.
But the model weights being available can give developers substantially more freedom than a purely API-based system.
For Tan, that freedom is strategically important.
What Is the “One Company” AI Nightmare?
Question: What does Tan mean when he warns about one company controlling AI?
Tan’s biggest concern is excessive concentration of AI capability in a single powerful proprietary company.
He described a scenario in which one organization has the strongest access to capital, the best AI researchers, and the most advanced systems, eventually becoming so dominant that competitors struggle to catch up.
He considers that a dangerous outcome.
The concern is not necessarily that a single company would intentionally behave badly.
The problem is structural.
If one provider controls the most capable AI models, developers and businesses may become dependent on that provider’s pricing, policies, access rules, safety decisions, and technical roadmap.
A healthy technology ecosystem generally benefits from competition because competing organizations can challenge each other’s assumptions and create alternative products.
Tan therefore wants a balance:
Frontier AI should remain commercially valuable enough to attract investment, while open-weight AI should remain strong enough to prevent excessive concentration.
That balance is central to his argument.
Could AI Model Distillation Help Smaller AI Labs Compete?
Question: Can distillation reduce the gap between small AI labs and frontier companies?
Potentially, yes. Distillation can give smaller developers a way to learn from the behavior and capabilities of more powerful systems without having to recreate every aspect of frontier-model development from scratch.
Building a frontier model requires enormous resources.
A leading AI lab may need:
- Large-scale computing infrastructure.
- Specialized AI researchers.
- Massive training datasets.
- Extensive model evaluation.
- Safety and alignment teams.
- Engineering infrastructure.
- Significant financial investment.
A smaller company may not be able to reproduce all of those resources.
Distillation can provide another route.
Instead of attempting to recreate a frontier model entirely independently, developers can potentially use its outputs to create a smaller, more focused system.
For example, a smaller model could be trained to perform specific tasks particularly well.
That could make AI development more accessible.
However, distillation does not magically transfer everything from a frontier model into a smaller one.
The quality of the resulting model depends on factors such as the data collected, the training process, model architecture, evaluation methods, and access conditions.
What Does This Mean for Frontier AI Companies?
Question: Why would frontier AI companies oppose unrestricted distillation?
Frontier companies have strong commercial reasons to protect their models from being replicated or commoditized.
Suppose a company invests billions of dollars developing a highly capable AI model. It then sells access through an API.
If another company can repeatedly query that model and use the responses to train a cheaper competing model, the original company could effectively help create its own competition.
That creates a difficult business incentive.
The frontier company wants customers to use its model because that generates revenue. But it may not want customers to use the same access to create a competing model.
This is why terms of service and technical controls have become important parts of the AI business.
Tan challenges that model.
His position is that companies should not have unlimited power to control what customers do with information generated by their services, particularly when the underlying AI was itself trained on broad collections of human knowledge.
The disagreement therefore involves both economics and policy.
Why the Copyright Argument Matters
Tan’s argument also connects the distillation debate to one of the biggest unresolved issues in AI: copyright.
AI companies have faced lawsuits and settlements concerning the use of copyrighted material in model training.
The TechCrunch report specifically points to Anthropic’s $1.5 billion copyright settlement, which had been approved in July 2026, as part of the broader context.
The argument is complicated because training and distillation are not identical activities.
Training a frontier model on books, websites, code, images, or other material involves one set of legal and technical questions.
Using outputs from a frontier model to train another model involves another.
But Tan sees a broader principle connecting them.
If AI companies can build powerful systems by processing enormous amounts of publicly accessible human knowledge, he questions whether the resulting intelligence should be treated as something that can be completely enclosed behind restrictive commercial rules.
That philosophical question could become increasingly important as AI-generated knowledge becomes more valuable.
What Are the Risks of Allowing More Distillation?
Question: Why might regulators still want restrictions on distillation?
The strongest argument for restrictions is that unrestricted distillation could make it easier to reproduce advanced AI capabilities, potentially creating security, safety, intellectual-property, and competitive risks.
Several concerns deserve attention.
Model Replication
A powerful model represents years of research and enormous investment. If competitors can cheaply reproduce important capabilities through API outputs, incentives to build frontier models could weaken.
Security
Anthropic’s allegations show why unauthorized access is a serious issue. Fraud and stolen credentials are fundamentally different from legitimate customer use and can create broader cybersecurity risks.
Safety
More capable models becoming easier to reproduce could increase the number of organizations able to deploy advanced AI systems.
That could make monitoring and safety enforcement more difficult.
Competitive Pressure
If frontier models can be rapidly distilled into cheaper alternatives, companies may struggle to recover the costs of developing them.
Regulatory Complexity
Rules would need to distinguish legitimate research, commercial competition, authorized distillation, unauthorized extraction, and outright credential theft.
That is not an easy line to draw.
Could the U.S. Create an “American Distillation Regime”?
Question: What could an American distillation regime look like?
Tan did not provide a detailed regulatory blueprint in the material reported by TechCrunch. His central idea is that American AI labs should be allowed to conduct legitimate distillation rather than having the government broadly prohibit the practice.
A potential policy framework would therefore need to answer several questions.
Who can perform distillation?
What counts as authorized access?
How should model providers disclose restrictions?
When does ordinary API usage become systematic model extraction?
What protections should exist for security-sensitive systems?
How should intellectual-property rights apply?
How should the rules distinguish American companies from foreign organizations?
These questions could become increasingly important as governments attempt to balance AI competition with security.
A blanket ban could potentially hurt smaller American AI companies.
No restrictions at all could potentially weaken the incentives for companies to invest in frontier models.
The difficult part is finding a middle ground.
What Does This Debate Mean for Open-Weight AI?
The debate could ultimately strengthen the case for open-weight development.
If developers have access to powerful open-weight models, they do not necessarily need to rely entirely on closed APIs to experiment with advanced AI.
Open-weight models can also create competition at multiple levels.
A company could build a specialized model for healthcare administration, coding, education, customer service, scientific research, or other applications without having to operate a frontier-scale laboratory.
That does not mean every open model will outperform proprietary systems.
Frontier companies are still likely to have advantages in compute, research talent, infrastructure, and access to capital.
But open-weight models can provide alternatives.
That is exactly what Tan wants.
His concern is that the AI industry could otherwise become dominated by a very small number of proprietary providers.
What Should AI Developers and Students Learn From This?
Question: Why should students and young AI professionals care about AI model distillation?
Because the future AI industry will not be defined only by who builds the largest model. It will also be shaped by how developers adapt, specialize, compress, evaluate, and deploy existing models.
For students and early-career professionals, the debate highlights several useful technical concepts.
First, understand how teacher-student model training works.
Second, learn how APIs can be used to build AI applications while respecting usage policies.
Third, understand the difference between open-weight and closed-weight models.
Fourth, learn about model evaluation because a smaller model is useful only if it maintains the capabilities required for its intended task.
Finally, understand the policy side of AI.
The technology is increasingly connected to copyright, cybersecurity, national competitiveness, and regulation.
AI professionals who understand both the engineering and policy dimensions will be better prepared for this changing landscape.
Why This Debate Could Shape the Next AI Era
The argument between Garry Tan and companies seeking tighter controls over distillation reflects a much larger transition in artificial intelligence.
During the early generative-AI boom, the primary question was often:
Who can build the most powerful model?
The next question may be:
Who gets to build on top of the most powerful models?
That distinction matters.
If only a handful of companies can develop advanced systems, AI could become increasingly centralized.
If smaller companies can freely access and learn from powerful models, innovation could spread more rapidly.
But unrestricted access could also create incentives for model replication and potentially increase security and safety risks.
The challenge is therefore not simply choosing “open” or “closed.”
It is designing an ecosystem in which frontier companies have enough economic incentive to keep investing while developers have enough freedom to innovate.
Tan believes the U.S. should lean more toward that second side of the equation.
FAQ: AI Model Distillation and Garry Tan
What is AI model distillation?
AI model distillation is a technique in which a smaller or newer AI model learns from the outputs or behavior of a more capable model. It can be used legitimately to create smaller, specialized, or more efficient AI systems.
What does Garry Tan want U.S. AI labs to do?
Garry Tan wants smaller U.S. open-weight AI labs to have the freedom to use legitimate distillation techniques on American frontier AI models. He argues that this could strengthen competition and create more open-weight alternatives.
Is Garry Tan supporting stolen credentials or hacking?
No. Tan’s argument, as reported by TechCrunch, is about allowing American companies to use legitimate access to frontier AI systems. Anthropic’s allegations concerning concealed identities, fraud, and stolen credentials describe a different category of alleged activity.
Why does Anthropic want tighter controls on distillation?
Anthropic has alleged that Chinese AI labs have conducted illicit distillation attacks and has called for U.S. regulators to address the issue. The company’s concern is that unauthorized extraction of model capabilities could undermine security and the competitive position of frontier AI developers.
Why are open-weight AI models important?
Open-weight models can give developers more control and flexibility than purely closed API-based systems, depending on their licenses. Supporters argue that they can increase competition, innovation, and access to AI capabilities.
Could distillation threaten frontier AI companies?
Potentially. If competitors can use legitimate access to a powerful model to develop cheaper competing systems, frontier AI companies could face pressure on their competitive advantage and revenue. This is one reason the rules surrounding model access and downstream use are becoming increasingly important.
The Bottom Line
Garry Tan’s argument is not simply about whether AI companies should be allowed to copy each other. It is about who should control the spread of AI capabilities. He wants U.S. open-weight labs to have legitimate access to distillation so they can compete with frontier AI companies and reduce the risk of excessive concentration.
Anthropic’s allegations about illicit distillation demonstrate why security and access controls matter, while Tan’s position highlights the opposite concern: overly restrictive rules could leave the AI ecosystem dominated by a handful of proprietary providers.
The likely long-term debate will therefore be less about whether distillation exists and more about where regulators draw the line between legitimate learning, commercial competition, intellectual property, and unauthorized extraction.
For anyone following AI in 2026, that line could help determine whether the next generation of AI is dominated by a few frontier labs,or powered by a much broader ecosystem of open-weight developers.
Want to understand where open AI, frontier models, and AI regulation are heading next? Keep following Kalinga.ai for more explainers on the technologies and policy decisions shaping the AI industry.