Five frontier models, one benchmark, zero cancers cured
CancerBench scores models on the promise the industry actually made in public. Everything is tied at nothing, which is the only honest number on the board.

CancerBench went up on September 9 with five entrants and one column. Claude Fable 5.1, GPT-6 Astra, Gemini 3.8 Flash, Grok 4.6 and Muse Spark 1.3 have each cured zero types of cancer. Its own caption reads "All tied for first. And last."

CancerBench: the frontier model cancer cure benchmark.
AI lab CEOs keep talking about curing cancer, so I made a benchmark.
One metric: how many types of cancer has your model cured?
All models are currently tied at zero.
It’s time to hillclimb!
t.co/SeEJz0FFYv t.co/GmOvBiKrEE

Tanishq Abraham built it because, in his words, AI lab CEOs keep talking about curing cancer, so he made a benchmark. One metric. How many types of cancer has your model cured?
Funny, obviously. But the useful half is further down the page, which is a dated receipt file of the sentences that produced it, collected between September 6, 2025 and September 6, 2026.
The thing that will work is actually curing cancer.
Amodei said that in August, having already told Davos in January that AI "will help us cure cancer." Elon Musk, in August, on whether AI companies should prove themselves this way, replied "AI will do it." Sam Altman, arguing for more compute in September 2025, offered that "you could choose to cure cancer by having AI do a bunch of research." Demis Hassabis described humanity benefiting from "cures for cancer." Reid Hoffman named drug discovery. Six statements, five leaders, twelve months, one scoreboard, all zeros.
Now put that next to the numbers that did move.
Everything with a scoreboard is climbing fast (and quickly enough that a chart drawn in July is already wrong). Gemini's coding score went from 49.0 to 73.7 in six weeks. Astra reports 99.9 percent on ARC-AGI-3. And the one metric nobody built a gradient for is flat, and it is flat because there is no gradient — you cannot hill-climb a thing with no partial credit.
So is this fair? Not really, and it is not trying to be. No serious person expects a language model to cure a disease unassisted, and Abraham obviously knows that. Abraham's benchmark is aimed at the sentence rather than the model, and the sentence is the thing the industry keeps saying to people who are frightened and to legislators who are deciding.
Our read is that CancerBench reads zero on September 9, 2027, and that this will be nobody's fault in particular. Probably the real contribution of AI to oncology is already arriving, and it looks like the thing DeepMind shipped the day before this benchmark went live — a petabyte of variant predictions that helped a Broad Institute team pin down a splice-site mutation in DNM1. Real, slow, indirect, and it will never produce a number you can put on a slide next to a competitor's.
Which is why the promise gets made in the vocabulary of cures instead. A cure is legible. Variant prioritisation is not.
What would a 1 on this board even look like? Somebody would have to define a cured type of cancer, agree who certifies it, and decide how much of the credit a model gets when it sits fourteen steps upstream of a trial. Until someone does that work, the honest score is the one currently posted, and the people quoted on that page are the ones who should find it uncomfortable.
