Skip to content
By AI

Anthropic Finally Answered the Open Weights Letter. A Chinese Lab Answered It Louder.

Amodei says Anthropic never wanted a ban. His real objection is irreversibility. Days earlier, a Chinese lab shipped 2.8 trillion open parameters.

Anthropic Finally Answered the Open Weights Letter. A Chinese Lab Answered It Louder., by Deepak Gupta on guptadeepak.com

Anthropic was the most conspicuous name missing from Nvidia's open weights letter, and for three days the silence read as strategy. On July 27, Dario Amodei published a response under his own byline with one sentence doing most of the work: "Anthropic has never advocated for a ban on open-weights models," with the emphasis his. Eleven days earlier, a Chinese lab had shipped the largest open weight model anyone has ever released. The argument and the evidence arrived in the same fortnight.

TL;DR

  • Amodei says Anthropic has never sought a ban and calls capable-but-safe open weight models a public good.
  • His actual disagreement is narrow: he disputes the claim that openness favours defenders over attackers, using bioweapons as the asymmetry case.
  • He proposes three measures instead of a ban: chip controls, a crackdown on industrial-scale distillation, and safety testing for sufficiently capable models regardless of whether weights ship.
  • Anthropic told the Senate Banking Committee in June that Alibaba's Qwen lab ran 28.8 million exchanges through roughly 25,000 fraudulent accounts over 44 days against Claude.
  • Moonshot's Kimi K3 launched July 16 at 2.8 trillion parameters, with full weights following on July 27. It activates 16 of 896 experts per token.

What Amodei actually argued

The position he published is narrower than the caricature, and narrower than I expected.

He agrees open weights expand access and competition. He calls open weight models without dangerous capabilities a public good, on the grounds that they cost nothing beyond compute and deliver real value to businesses, developers, and researchers. That is not a hedge. It is most of the Nvidia letter's premise.

Where he breaks from it is on a single asymmetry. The letter implies openness helps defenders more than attackers. Amodei disputes that, and his example is biological weapons: a capable enough model could help someone weaponise a pathogen using materials that are already obtainable, on a timeline of weeks, while building real defences against that threat takes years. Once weights are public, there is no recall if a capability nobody tested for surfaces later.

You do not have to accept the bioweapons case to see the structural point. It is an argument about irreversibility, not about openness. Every other lever in this debate can be adjusted after the fact. Published weights cannot.

Instead of a ban, he asks for three things:

MeasureTargetWho it constrains
Chip and equipment export controlsAuthoritarian governmentsNation states, not developers
Crackdown on industrial-scale distillationSystematic capability extractionLabs running extraction campaigns
Safety testing for sufficiently capable modelsUntested frontier capabilityOpen and closed labs alike

The third one is the one worth noticing, because it costs Anthropic something. Mandatory testing for sufficiently capable models applies to closed labs too, and Anthropic is a closed lab. He has also floated an international body to run that testing, which is a heavier lift than anything the Nvidia letter proposes.

He was also honest about the enforcement problem, which most policy writing skips. Accounts running a distillation campaign are usually only identifiable after a good deal of the damage is done. That is a real argument for targeted mechanisms over trusting the market, and it is also an admission that his preferred remedy is hard to operate.

The distillation numbers, since everyone rounds them

The specifics matter here and the coverage keeps blurring them. In a letter to senior members of the Senate Banking Committee dated June 10, 2026, Anthropic alleged that operators connected to Alibaba ran 28.8 million exchanges with Claude through approximately 25,000 fraudulent accounts across a 44-day window between April 22 and June 5. Anthropic described it as the largest known distillation campaign ever run against a commercial model, aimed at lifting Qwen toward the performance of Anthropic's frontier Mythos Preview on software engineering, multi-step reasoning, and cybersecurity.

Anthropic's separation of open weights from distillation is the part of its position closest to what several of the Nvidia letter's own signatories want. The letter argues the two are different problems and regulating one will not fix the other. Amodei's objection is not that they are the same problem. It is that treating them as fully independent lets the distillation question quietly drop off the agenda.

Meanwhile the model at the centre of the argument got bigger

While the policy fight was running, Moonshot AI answered a different question: who holds the record for the largest open weight model. Kimi K3 launched on July 16 at 2.8 trillion parameters, with the full weights following on July 27.

The scale number is the headline. The architecture is what makes the scale usable.

PropertyKimi K3Why it matters
Total parameters2.8 trillionLargest open weight model released to date
Experts896, with 16 active per tokenRoughly 1.8% of the pool fires on any given token
Context window1 million tokensLong-document and repository-scale work
ModalityText and images nativelyNo bolted-on vision adapter
Weights on disk~1.4 TB, even at 4-bitOpen does not mean small
Recommended hardwareSupernode configs, 64+ acceleratorsSelf-hosting is out of reach for most teams

Mixture-of-experts is the whole trick. You get something near the knowledge capacity of an enormous model while paying inference costs closer to a much smaller one. Moonshot also credits two in-house techniques, Kimi Delta Attention and Attention Residuals, for better scaling efficiency than K2. If you want the background on why parameter count and active parameters are different numbers, I covered that distinction in LLM vs SLM.

On results: in blind testing by the evaluator Arena, developers reportedly preferred K3 over leading closed US models on front-end coding specifically. On Artificial Analysis's composite leaderboard it posted a large jump over its predecessor and landed just behind Claude, ahead of prior-generation models from both major US labs while still trailing the current frontier overall. Note the shape of that: a narrow win on one task category, a respectable composite, not parity.

And note the practical ceiling. At 1.4 TB and 64 accelerators, "open weights" here means auditable and forkable, not runnable on your own hardware. For almost every team, hitting Moonshot's API remains the realistic path. Openness at this scale is a licensing property more than a deployment one.

Why the timing sharpens the argument

Moonshot founder Yang Zhilin has been explicit that growth through openness is the strategy, in a way the closed US labs have not matched. Since K3 shipped, Moonshot's daily revenue has reportedly grown at least sixfold, and the release knocked competitor stocks lower, with Z.ai and MiniMax both taking hits and Alibaba dipping.

Set that beside the policy fight and the picture gets sharper than either story alone. A coalition of major American companies is lobbying against broad restrictions on open weight models. Anthropic argues the real risks are narrower and specific: chip access, distillation, and untested capability. And in the middle of that argument, a Chinese lab shipped the largest open weight model on record, one already winning developer preference on a specific coding category.

None of this resolves the policy question. What it does is remove the option of treating the question as theoretical. The capability gap that the whole debate is ultimately about is closing in public while the rules are still being drafted. I wrote about the commercial version of this shift in how AI-native pricing is changing, and the security version of it is the same story: the thing being regulated keeps moving faster than the regulation.

For anyone building on these models rather than arguing about them, the practical read is narrower. Capability is now genuinely multi-sourced, cost pressure from open weight models is real and increasing, and any architecture that assumes a single frontier vendor is taking on a risk it probably has not priced.

Frequently Asked Questions

Does Anthropic want open weight models banned?

No. Dario Amodei stated on July 27, 2026 that "Anthropic has never advocated for a ban on open-weights models," and described open weight models without dangerous capabilities as a public good. His disagreement with the Nvidia letter is about whether openness favours defenders over attackers, not about openness itself.

What does Anthropic want instead?

Three measures: tighter export controls on advanced chips and chipmaking equipment flowing to authoritarian governments, enforcement against industrial-scale distillation, and mandatory safety testing for sufficiently capable models whether or not their weights are published. He has also raised the idea of an international body to run that testing.

What is the Alibaba distillation allegation?

In a June 10, 2026 letter to the Senate Banking Committee, Anthropic alleged that operators connected to Alibaba's Qwen lab ran 28.8 million exchanges with Claude through roughly 25,000 fraudulent accounts over a 44-day window from April 22 to June 5, 2026, in order to extract training signal. Anthropic called it the largest known distillation campaign against a commercial model.

How big is Kimi K3 and when was it released?

Kimi K3 launched on July 16, 2026 with 2.8 trillion total parameters, and Moonshot released the full open weights on July 27. It is a mixture-of-experts model with 896 experts that activates 16 per token, supports a one million token context window, and handles text and images natively.

Can I actually run Kimi K3 myself?

Realistically, no. The weights are about 1.4 TB even in a 4-bit format, and Moonshot recommends supernode configurations with at least 64 accelerators. For most teams the open weights mean the model is auditable and forkable rather than self-hostable, and the API remains the practical path.

Is Kimi K3 better than Claude or GPT?

On one category, reportedly. Developers in blind Arena testing preferred K3 over leading closed US models on front-end coding tasks. On Artificial Analysis's composite leaderboard it landed just behind Claude, ahead of prior-generation US models but still short of the current frontier overall. A category win is not parity.

What does distillation actually mean here?

Not copying weights. A smaller student model is trained on huge volumes of a larger teacher model's outputs until it reproduces a meaningful share of the teacher's behaviour, without ever touching the teacher's internal parameters. It reconstructs behaviour from outputs, which is why it sits in a genuinely murky legal category compared with straightforward code theft.

Get the newsletter

New writing on identity, AI security, and building software, delivered when it ships. No tracking pixels, no funnels, unsubscribe with one click.