A radiologist examines mammogram scans on a diagnostic workstation screen
Technology & Solutions

AI-supported breast screening found 29 percent more cancers

Adobe Stock

In the first randomised trial of artificial intelligence inside a national screening programme, 105,934 Swedish women were assigned either to AI-supported reading or to the standard two radiologists. The AI-supported group turned up 6.4 cancers for every 1,000 women screened, against 5.0, and needed 44 percent less reading to do it.

Peer-reviewed Checked 28 July 2026
Key numbers
6.4
Cancers found for every 1,000 women screened
up from 5.0 with two radiologists and no AI
61,248
Screen readings needed across the AI-supported group
down from 109,692 in the group read the standard way, a cut of 44.2 percent
80.5 percent
Share of the cancers present that the screening caught
up from 73.8 percent
1.55
Cancers appearing between screening rounds, per 1,000 women
down from 1.76, a difference the trial could not separate from chance

In a European breast screening programme every mammogram is read twice, by two radiologists working independently, a practice known as double reading. One pair of eyes misses things. In four towns in southwest Sweden, 105,934 women were randomly assigned either to that standard or to a version of it with artificial intelligence added, software that sorted each set of images and marked what looked suspicious. Among the 53,043 women screened with AI support, radiologists found 338 cancers. Among the 52,872 read the usual way, they found 262.

How we know

The trial is called MASAI, and it ran inside Sweden's national screening programme at Malmö, Lund, Landskrona and Trelleborg between April 2021 and December 2022. Women were allocated one to one. The software did two jobs: it sorted each examination towards a single reader or a double one, and it highlighted suspicious findings for the radiologist looking at the images. The screening results, published in The Lancet Digital Health in March 2025, put cancer detection at 6.4 per 1,000 women screened against 5.0, a 29 percent increase the analysis puts well outside chance. Recalls for further tests rose slightly, by an amount too small to separate from chance, and the false alarm rate did not move. The reading itself nearly halved: 61,248 screen readings in the AI-supported group against 109,692 in the control, a cut of 44.2 percent. The final results, in The Lancet in January 2026, tested the harder question. The trial's primary measure was the interval cancer rate, the cancers that surface between one screening round and the next after a screen has missed them. That rate was 1.55 per 1,000 women with AI support and 1.76 without, a difference inside the margin the trial had set for showing AI was no worse, but not large enough to show it was better. Sensitivity, the share of the cancers present that the screen actually caught, rose from 73.8 to 80.5 percent. Specificity, the share of healthy women correctly cleared, was 98.5 percent in both groups.

Why it matters

Double reading is expensive in the one resource screening services are short of, which is breast radiologists. Cutting the reading load by 44 percent without losing accuracy hands a stretched service half its reading time back, and in this trial it found more cancer while doing it: 58 more small tumours of the stage doctors call T1, and 46 more that had not yet reached the lymph nodes. That is the shape of a good result in screening, because a small contained tumour is more treatable than a large one. What the trial cannot yet say is whether any of this saves lives. It counted cancers, not deaths. The interval cancer signal points the right way, with fewer of those cancers invasive, 75 against 89, and fewer of them large or aggressive, but the comparison was not statistically significant and the authors claim only that AI support is not worse on it. The research question has moved. It is no longer whether software can read a mammogram, but whether a screening service should let it.

What is not solved yet

The trial measured cancers found, not lives saved. Its primary measure, the rate of cancers appearing between screening rounds, came out no worse rather than better: 1.55 per 1,000 women against 1.76, a gap well inside the range chance could produce. Whether catching more cancer earlier means fewer women die of breast cancer is a question this trial was not built to answer. The findings also rest on one country, four screening sites, one type of mammography machine and one AI system, read by radiologists with moderate to high experience, and the trial recorded no data on race or ethnicity, so how the result transfers to other populations, other equipment and other software is untested.

Common questions
Does this mean AI is replacing radiologists?

No. Every set of images in the AI-supported group was still read by at least one radiologist. The software did two jobs, sorting each examination towards a single reading or a double one and marking suspicious findings for the doctor reading them. What fell was the number of readings, from 109,692 to 61,248, not the presence of a human.

Does it save lives?

Not shown. The trial counted cancers, not deaths, and its primary measure, the rate of cancers appearing between screening rounds, came out no worse rather than better, at 1.55 per 1,000 women against 1.76. Finding cancers earlier and smaller is a reasonable step towards better survival, but this trial did not measure survival.

Did it cause more false alarms?

No. The false positive rate was effectively unchanged between the two groups, and recalls for further tests rose only slightly, by an amount the analysis could not distinguish from chance. The share of recalls that turned out to be cancer was in fact higher in the AI-supported group.

Did the AI flag things that would never have caused harm?

It found more in situ cancers, the non-invasive kind confined to the milk ducts, 68 against 45. About half of that extra detection was high grade, the more concerning end of the range, and there was no increase at all in the lowest grade. That pattern argues against the software simply picking up harmless changes, though the trial did not follow women long enough to measure long-term outcomes.