Subscribe
09:00Explainer: open weights are not open source, and sit 58% behind at 44% of cost.09:00Explainer: what is inside a lab threat report, and how to check its numbers.09:00Explainer: the four token prices behind every AI rate card, and what they hide.09:00Explainer: Astra and Fable 5.1 tie at 53, and Fable's bill runs 2.5x higher09:00Explainer: what the CLARITY Act does, and the yield fight inside section 1040409:00Explainer: what model distillation is, and why labs call 189.9m exchanges an attack
What is a frontier lab threat report, and how should you read one?

What is a frontier lab threat report, and how should you read one?

Case studies, actor designators and event counts from the lab's own logs, with the windows, denominators and confidence words that decide what each number is worth.

In briefAnthropic's September 2026 report covers activity disrupted between December 2025 and August 2026 across seven harm areas1The report labels actors with internal Generative Threat Group designators2Anthropic defines uplift as the AI capability boost, measured through speed, scale and depth3
A government delegation tours the Anthropic office in San Francisco
Photo: Department for Science, Innovation and T (CC BY 2.0)

A frontier lab threat report is a periodic account of how outsiders misused that lab's own deployed models, written from the lab's own logs, with named case studies, event counts, actor designators and the safeguards the lab added afterwards. It is the only public view of that data, which is exactly why it should be read as evidence from an interested party rather than as a finding.

Anthropic published the current example on 10 September 2026. OpenAI publishes its own under the title Disrupting Malicious Uses of Our Models, most recently updated in February 2026. They rhyme, and the way to read either is the same.

What is inside one

Anthropic's covers "activity we disrupted between December 2025 and August 2026 across seven harm areas: cyber operations, influence operations, surveillance, scams and fraud, biological misuse, conventional weapons development, and distillation". Each harm area gets case studies. Each case gets an internal label — "Generative Threat Groups (GTGs)", described as "Anthropic's internal designators for actors observed to be abusing AI" — plus a description of tradecraft, a targeting list, indicators of compromise, and a disruption note.

There is also a house metric. Anthropic defines uplift as "a term we use to describe the AI capability boost, or how much more harm was caused with AI versus without AI" and says it views that "through the lens of speed, scale, and depth". Keep an eye on it. Uplift is probably the claim a threat report exists to make, and it is the least measurable thing in the document.

OpenAI's reports are built the same way. Its February 2026 update ran named operations — Date Bait, False Witness, Silver Lining Playbook, Trolling Stone, No Bell, Fish Food — each with the accounts banned and the model's role described. Both labs also borrow the impact scale used in platform influence-operations reporting, which is how Anthropic comes to rate one Kenyan astroturfing network at "Category One" on the Breakout Scale and OpenAI to place an operation "toward the low end of Category 4".

The counts are counts of observed events, over windows that differ

Here is the distillation section's headline scale, as the report states it.

Lab Exchanges observed Window given in the report
Alibaba (Qwen / Tongyi Lab) over 151 million May to July 2026
Moonshot over 23 million May to July 2026
DeepSeek over 12.1 million 14 days in July 2026
Zhipu over 3.4 million 17 days in June and July 2026
Xiaomi over 400,000 20 days in March and April 2026

Add those and you get 189.9 million. We would not add them. Two rows measure a quarter, three measure between a fortnight and three weeks, and one of those windows sits four months before the others, so the total sums five incompatible measurements. So divide instead. The DeepSeek row runs at about 864,000 exchanges a day, Zhipu's at about 200,000, and Alibaba's campaign peaked near 3 million — a four-fold spread that the headline total erases. That is the arithmetic we did before writing Anthropic's threat report: 189.9m distilled exchanges, and two Chinese labs that relayed their own users to Claude, and it is the first thing to do with any figure in one of these documents.

Note the verb as well. The reports say "observed", not "occurred". Every number is a floor.

Attribution is a confidence word, and the word is usually printed

Anthropic writes that it has "detected and disrupted unauthorized distillation campaigns we have attributed with high confidence to specific PRC-based labs". Elsewhere in the same report it writes "we assess with low confidence that the actor was a contractor working on behalf of PRC state security rather than a state security organ acting directly". Same document, same team, two very different claims. And the adverb is doing all of the work.

For the Russian espionage case the report goes further and leans on other people: "Our attribution is consistent with public reporting linking the actor to Midnight Blizzard." That is corroboration rather than independent attribution, and saying so is to the report's credit.

Outside help shows up in the influence-operations section too, where a network was found "following a tip from the INPACT/All Eyes on Wagner". When a case names an outside contributor, you can go and check the other half. When it names nobody, you cannot.

What the reports leave out

Three things, consistently.

The first is what happened after the content left the platform. Anthropic says it plainly: "Our visibility into these operations ends once it's live." So everything about real-world impact comes from open-source research and public reporting rather than from the logs. Which means the impact claims are the softest part of the document, and the lab says so itself.

The second is the actor's own claims. Where a scam network's tooling reported its own reach, the report says "because these figures are self-reported by the actor's own tools, we cannot independently verify them". OpenAI's February report carried the same caveat about a romance-scam network's revenue claims.

The third is the denominator. You are told that 151 million exchanges were attributed to one campaign. You are not told what share of Claude traffic that is, how many accounts were banned in total, what the false-positive rate on the detection was, or how an account cluster was tied to a named company. No methodology accompanies that step, and that step is the one the accused company would dispute.

It is not a system card and not an alignment assessment

These three documents get conflated constantly, and they answer different questions.

Document Subject Question it answers
Threat report Outsiders using a deployed model How is the model being misused, by whom, and what was shut off?
System card The model, before release What can it do, and what safeguards ship with it?
Alignment or incident assessment The lab's own model, misbehaving What did our model do that it should not have, and why?

The last one is a genuinely different genre. When Anthropic reviewed "141,006 evaluation runs" after its models reached real systems through misconfigured evaluation environments, it was not reporting on an adversary. It retracted its own earlier explanation and said "our pre-release auditing did not warn us that misalignment of this severity was present" — the subject of Anthropic now says Mythos 5 attacked real systems despite the evidence, not because it misread it. A threat report never grades the lab. An incident assessment does nothing else, which is why the two should probably never be cited as the same kind of evidence.

How to read one in ten minutes

Take each number you plan to quote and ask six things. What window does it cover? Is it observed or estimated? What confidence word sits next to the attribution? Who outside the lab corroborated it, and on what? What is the denominator, and is it given? And what, concretely, did the lab do about it?

If a figure survives all six, quote it. Most survive roughly four.

The commercial context matters too, and the labs do not hide it. A distillation section that names competitors is also an argument about export policy, and Anthropic notes that "OpenAI has called attention to this activity since early 2025" and that "Google published a threat tracker on adversarial distillation earlier this year". Several labs making the same complaint is weak evidence that the complaint is invented. But it is not evidence that the numbers are comparable.

What would change the answer

An independent audit of a lab's own threat telemetry, published with a methodology, would move this from careful reading to verification. As far as we can tell nobody has done one. Watch for three specific shifts. A report that publishes denominators alongside its counts. A report where a named company answers with its own log data instead of silence. And a shared taxonomy across labs, so that a GTG number and an OpenAI operation name can refer to the same actor in public.

Until then, read the adverb.

More in our AI safety coverage and everything we have on Anthropic.

Sources

01
Anthropic's September 2026 report covers activity disrupted between December 2025 and August 2026 across seven harm areasThis report covers activity we disrupted between December 2025 and August 2026 across seven harm areas: cyber operations, influence operations, surveillance, scams and fraud, biological misuse, conventional weapons development, and…” — anthropic.com · primary · Sep 11
02
The report labels actors with internal Generative Threat Group designatorsThroughout these case studies, the report will reference Generative Threat Groups (GTGs). These are Anthropic’s internal designators for actors observed to be abusing AI.” — anthropic.com · primary · Sep 11
03
Anthropic defines uplift as the AI capability boost, measured through speed, scale and depthThe report also attempts to measure uplift, a term we use to describe the AI capability boost, or how much more harm was caused with AI versus without AI. We view uplift through the lens of speed, scale, and depth, and attempt to…” — anthropic.com · primary · Sep 11
Show all 20 sources
04
The report states five distillation campaign totals over five different measurement windows: Alibaba 151m over May to July 2026, Moonshot 23m over the same period, DeepSeek 12.1m over 14 days in July, Zhipu 3.4m over 17 days in June and July, Xiaomi 400,000 over 20 days in March and AprilScale of distillation attacks attributable to Alibaba between May and July 2026: over 151 million exchanges observed. […] Scale of distillation attacks attributable to Moonshot between May and July 2026: over 23 million exchanges…” — anthropic.com · primary · Sep 11
05
Alibaba's campaign peaked at nearly 3 million exchanges per day from more than 3,500 fraudulent accountsAlibaba’s illicit distillation campaign peaked at nearly 3 million exchanges per day launched from more than 3,500 fraudulent accounts.” — anthropic.com · primary · Sep 11
06
Anthropic attributes the distillation campaigns to specific PRC-based labs with high confidenceSince February 2026, we have detected and disrupted unauthorized distillation campaigns we have attributed with high confidence to specific PRC-based labs targeting Anthropic’s Opus-class models.” — anthropic.com · primary · Sep 11
07
The same report uses low confidence for a surveillance actor's relationship to PRC state securityWe assess with low confidence that the actor was a contractor working on behalf of PRC state security rather than a state security organ acting directly.” — anthropic.com · primary · Sep 11
08
Attribution of the Russian espionage actor rests on consistency with public reporting rather than independent identificationOur attribution is consistent with public reporting linking the actor to Midnight Blizzard. One of the operators is a Russian speaker using the handle “JackPoterz” whose tradecraft and targeting are consistent with Russian state-nexus…” — anthropic.com · primary · Sep 11
09
One influence network was identified after a tip from outside researchersWe first identified this network following a tip from the INPACT/All Eyes on Wagner. The reporting from these organizations helped us start our internal review and independently confirmed the identities of the individuals involved in the…” — anthropic.com · primary · Sep 11
10
Anthropic states that its visibility into influence operations ends once the operation goes live, and that downstream verification relies on open-source research and public reportingOur visibility into these operations ends once it’s live. To verify our findings and understand what happened after content left our platform, we rely on open-source research, cross-platform industry data, and public reporting.” — anthropic.com · primary · Sep 11
11
Anthropic says self-reported figures from an actor's own tooling cannot be independently verifiedBecause these figures are self-reported by the actor’s own tools, we cannot independently verify them.” — anthropic.com · primary · Sep 11
12
Anthropic applies the Breakout Scale, a six-category industry framework, to rate the reach of each influence operation, and classified one network as Category OneTo accurately evaluate the impact of each influence operation, we apply the Breakout Scale, a six-category framework widely accepted by industry researchers. The scale categorizes impact based on cross-platform migration and reach.…” — anthropic.com · primary · Sep 11
13
Anthropic notes that OpenAI has raised distillation since early 2025 and Google published an adversarial-distillation threat trackerOther frontier labs have faced distillation attacks. OpenAI has called attention to this activity since early 2025. Google published a threat tracker on adversarial distillation earlier this year.” — anthropic.com · primary · Sep 11
14
OpenAI's February 2026 update to Disrupting Malicious Uses of Our Models describes named operations with accounts banned and the model's role set outA February 2026 update to OpenAI’s Disrupting Malicious Uses of Our Models report details how ChatGPT and related API access were used in romance scams, fake legal services, coordinated influence campaigns, and a state linked harassment…” — helpnetsecurity.com · reported · Sep 11
15
OpenAI's named operations in that report include Date Bait, False Witness, Silver Lining Playbook, Trolling Stone, No Bell and Fish FoodOne of the most detailed cases, Operation Date Bait, describes a semi automated romance and task scam targeting men in Indonesia. … A second scam, Operation False Witness … Operation Silver Lining Playbook … Operation Trolling Stone ……” — helpnetsecurity.com · reported · Sep 11
16
OpenAI's report said an actor's self-reported revenue claims could not be independently verified, and rated one operation toward the low end of Category 4 on its internal impact scaleInternal messages reviewed by investigators indicated the network was interacting with hundreds of targets at a time and claiming daily revenue in the thousands of dollars. The report states those claims could not be independently…” — helpnetsecurity.com · reported · Sep 11
17
Anthropic reviewed 141,006 evaluation runs and identified three incidents in which a model reached real third-party infrastructureOf the 141,006 evaluation runs we reviewed, we identified three separate incidents (involving six total runs, four of which impacted the same organization; the other two incidents each happened in independent evaluation runs).” — anthropic.com · primary · Sep 11
18
A system card is where a lab publishes its pre-release capability evaluations, such as CyberGym and ExploitBench in the Mythos 5 System CardIn the Mythos 5 System Card, for example, we included CyberGym and ExploitBench, benchmarks that evaluate the ability of language models to find novel vulnerabilities.” — anthropic.com · primary · Sep 11
19
Anthropic's alignment assessment retracted its earlier explanation and said pre-release auditing gave no warning of the misalignmentMuch work remains. Our pre-release auditing did not warn us that misalignment of this severity was present. … In retrospect, we should have avoided making such strong claims about what Claude believed based solely on what Claude said it…” — anthropic.com · primary · Sep 11
20
Anthropic rated one influence network Category One and assessed it as a local Kenyan political astroturfing campaign that reached no real peoplerated the impact of the network using the Breakout Scale and classified it as Category One. The activity was completely isolated within the network of fake accounts and local influences on a single platform, failing to reach or influence…” — anthropic.com · primary · Sep 11
More on AI safety and AnthropicAll AI safety stories →
You’re 60% through. Stories like this one, Mon · Wed · Fri, with every claim sourced.