What is a frontier lab threat report, and how should you read one?
Case studies, actor designators and event counts from the lab's own logs, with the windows, denominators and confidence words that decide what each number is worth.

A frontier lab threat report is a periodic account of how outsiders misused that lab's own deployed models, written from the lab's own logs, with named case studies, event counts, actor designators and the safeguards the lab added afterwards. It is the only public view of that data, which is exactly why it should be read as evidence from an interested party rather than as a finding.
Anthropic published the current example on 10 September 2026. OpenAI publishes its own under the title Disrupting Malicious Uses of Our Models, most recently updated in February 2026. They rhyme, and the way to read either is the same.
What is inside one
Anthropic's covers "activity we disrupted between December 2025 and August 2026 across seven harm areas: cyber operations, influence operations, surveillance, scams and fraud, biological misuse, conventional weapons development, and distillation". Each harm area gets case studies. Each case gets an internal label — "Generative Threat Groups (GTGs)", described as "Anthropic's internal designators for actors observed to be abusing AI" — plus a description of tradecraft, a targeting list, indicators of compromise, and a disruption note.
There is also a house metric. Anthropic defines uplift as "a term we use to describe the AI capability boost, or how much more harm was caused with AI versus without AI" and says it views that "through the lens of speed, scale, and depth". Keep an eye on it. Uplift is probably the claim a threat report exists to make, and it is the least measurable thing in the document.
OpenAI's reports are built the same way. Its February 2026 update ran named operations — Date Bait, False Witness, Silver Lining Playbook, Trolling Stone, No Bell, Fish Food — each with the accounts banned and the model's role described. Both labs also borrow the impact scale used in platform influence-operations reporting, which is how Anthropic comes to rate one Kenyan astroturfing network at "Category One" on the Breakout Scale and OpenAI to place an operation "toward the low end of Category 4".
The counts are counts of observed events, over windows that differ
Here is the distillation section's headline scale, as the report states it.
| Lab | Exchanges observed | Window given in the report |
|---|---|---|
| Alibaba (Qwen / Tongyi Lab) | over 151 million | May to July 2026 |
| Moonshot | over 23 million | May to July 2026 |
| DeepSeek | over 12.1 million | 14 days in July 2026 |
| Zhipu | over 3.4 million | 17 days in June and July 2026 |
| Xiaomi | over 400,000 | 20 days in March and April 2026 |
Add those and you get 189.9 million. We would not add them. Two rows measure a quarter, three measure between a fortnight and three weeks, and one of those windows sits four months before the others, so the total sums five incompatible measurements. So divide instead. The DeepSeek row runs at about 864,000 exchanges a day, Zhipu's at about 200,000, and Alibaba's campaign peaked near 3 million — a four-fold spread that the headline total erases. That is the arithmetic we did before writing Anthropic's threat report: 189.9m distilled exchanges, and two Chinese labs that relayed their own users to Claude, and it is the first thing to do with any figure in one of these documents.
Note the verb as well. The reports say "observed", not "occurred". Every number is a floor.
Attribution is a confidence word, and the word is usually printed
Anthropic writes that it has "detected and disrupted unauthorized distillation campaigns we have attributed with high confidence to specific PRC-based labs". Elsewhere in the same report it writes "we assess with low confidence that the actor was a contractor working on behalf of PRC state security rather than a state security organ acting directly". Same document, same team, two very different claims. And the adverb is doing all of the work.
For the Russian espionage case the report goes further and leans on other people: "Our attribution is consistent with public reporting linking the actor to Midnight Blizzard." That is corroboration rather than independent attribution, and saying so is to the report's credit.
Outside help shows up in the influence-operations section too, where a network was found "following a tip from the INPACT/All Eyes on Wagner". When a case names an outside contributor, you can go and check the other half. When it names nobody, you cannot.
What the reports leave out
Three things, consistently.
The first is what happened after the content left the platform. Anthropic says it plainly: "Our visibility into these operations ends once it's live." So everything about real-world impact comes from open-source research and public reporting rather than from the logs. Which means the impact claims are the softest part of the document, and the lab says so itself.
The second is the actor's own claims. Where a scam network's tooling reported its own reach, the report says "because these figures are self-reported by the actor's own tools, we cannot independently verify them". OpenAI's February report carried the same caveat about a romance-scam network's revenue claims.
The third is the denominator. You are told that 151 million exchanges were attributed to one campaign. You are not told what share of Claude traffic that is, how many accounts were banned in total, what the false-positive rate on the detection was, or how an account cluster was tied to a named company. No methodology accompanies that step, and that step is the one the accused company would dispute.
It is not a system card and not an alignment assessment
These three documents get conflated constantly, and they answer different questions.
| Document | Subject | Question it answers |
|---|---|---|
| Threat report | Outsiders using a deployed model | How is the model being misused, by whom, and what was shut off? |
| System card | The model, before release | What can it do, and what safeguards ship with it? |
| Alignment or incident assessment | The lab's own model, misbehaving | What did our model do that it should not have, and why? |
The last one is a genuinely different genre. When Anthropic reviewed "141,006 evaluation runs" after its models reached real systems through misconfigured evaluation environments, it was not reporting on an adversary. It retracted its own earlier explanation and said "our pre-release auditing did not warn us that misalignment of this severity was present" — the subject of Anthropic now says Mythos 5 attacked real systems despite the evidence, not because it misread it. A threat report never grades the lab. An incident assessment does nothing else, which is why the two should probably never be cited as the same kind of evidence.
How to read one in ten minutes
Take each number you plan to quote and ask six things. What window does it cover? Is it observed or estimated? What confidence word sits next to the attribution? Who outside the lab corroborated it, and on what? What is the denominator, and is it given? And what, concretely, did the lab do about it?
If a figure survives all six, quote it. Most survive roughly four.
The commercial context matters too, and the labs do not hide it. A distillation section that names competitors is also an argument about export policy, and Anthropic notes that "OpenAI has called attention to this activity since early 2025" and that "Google published a threat tracker on adversarial distillation earlier this year". Several labs making the same complaint is weak evidence that the complaint is invented. But it is not evidence that the numbers are comparable.
What would change the answer
An independent audit of a lab's own threat telemetry, published with a methodology, would move this from careful reading to verification. As far as we can tell nobody has done one. Watch for three specific shifts. A report that publishes denominators alongside its counts. A report where a named company answers with its own log data instead of silence. And a shared taxonomy across labs, so that a GTG number and an OpenAI operation name can refer to the same actor in public.
Until then, read the adverb.
More in our AI safety coverage and everything we have on Anthropic.



