Evidence-first health & aging science Newsletter
Live Well News Live Well News
Back to Longevity & Aging

News · Longevity & Aging

An AI system found 90% of esophageal cancers on routine chest CT scans but just 52.5% of precancerous lesions

Built by Alibaba's research arm and tested on more than 80,000 people, the tool wrongly flagged about 15 of every 1,000 people without the disease. It was trained in China, where the common tumor type differs from that in the West.

A clinician in a face mask, glasses and blue gloves holds up a chest X-ray film to examine it.
Summary
  • An AI system read ordinary chest CT scans for signs of esophageal cancer in more than 80,000 people.
  • In tests at eight centers it detected 90.0% of cancers and 52.5% of precancerous lesions.
  • About 15 of every 1,000 people without the disease were wrongly flagged.
  • With the system's help, the average detection rate of 17 radiologists rose from 71.9% to 85.7%.
  • It was trained only in China, and the study did not test whether using it reduces deaths.

A chest scan ordered to look at the lungs also photographs the esophagus, the tube that carries food to the stomach. Radiologists have long regarded that picture as too poor to show early cancer there. A study published in Nature Medicine in September reports that software can often do it, on the plainest kind of computed tomography (CT) scan, while wrongly flagging about 15 of every 1,000 people who do not have the disease.

The system picked out 90.0% of esophageal cancers in a test on more than 11,000 patients. It did less well on the earliest disease, catching about half of precancerous lesions, the abnormal patches that can turn into cancer. And it was built and mostly tested in China, where esophageal cancer is a different disease from the one most common in Europe and North America.

Why esophageal cancer is hard to catch early

Esophageal cancer usually announces itself late. An estimated 511,000 new cases were diagnosed worldwide in 2022, and about 445,000 people died from the disease, a ratio that reflects how often it is found at an advanced stage.

The test that finds it early is endoscopy, in which a camera on a flexible tube is passed down the throat so that doctors can look at the esophagus and take a tissue sample. It is invasive and difficult to use as a screening tool in large populations. The paper names the same problem: the absence of accurate, noninvasive, scalable screening tools.

China has tried endoscopic screening at scale in regions where the disease is common. A cohort study conducted in six areas in China from 2005 to 2015 followed 637,500 residents. Deaths from cancers of the esophagus and stomach were 57% lower among people who were screened than among those never invited. Uptake was the weak point. Of 338,017 people invited, 113,340, about a third, were screened. That comparison was not randomized, and people who accept screening tend to differ from those who decline.

A CT scan is far easier to get than an endoscopy, and millions are already taken for other reasons. The obstacle has been the organ. In the authors’ description, the esophagus is a hollow tubular structure prone to collapse and motion artifacts, making small early malignant lesions difficult to distinguish from normal tissue. Motion artifacts are blurring caused by movement, and malignant means cancerous. They call finding cancer on such scans a task historically considered impossible.

How the AI system reads a chest CT

The system is called EAGLE, and it comes from DAMO Academy, the research arm of the Chinese technology company Alibaba, working with hospitals. The company describes it as a new AI model designed for the early detection of esophageal cancer. It joins earlier models from that group for cancers of the pancreas, stomach, bowel and liver.

It works on noncontrast scans, meaning those taken without the injected dye that makes blood vessels and tumors stand out. The system works in two stages. First it locates the esophagus within the three-dimensional CT scan, then analyzes it, searches for lesions and calculates the probability that a finding is malignant. It also marks where the suspicious spot is.

The scale of the testing is what sets the study apart. The model was trained on 6,813 patients from two centers and validated across 12 centers in three countries involving 80,612 patients. The three countries were China, the Czech Republic and Australia. The tests covered two situations, described as opportunistic and population-based screening settings. Opportunistic screening means checking a scan that was taken for some other purpose. Population-based screening means inviting a whole group of people to be tested.

What the screening tests showed

Two terms carry the results. Sensitivity is the share of people with the disease whom the test correctly flags. Specificity is the share of people without it whom the test correctly clears.

TestPatientsResult
Existing scans, eight centers11,46690.0% of cancers and 52.5% of precancerous lesions detected; 98.5% specificity
Low-dose lung screening scans, two centers1,60788.4% sensitivity; 99% specificity
Hospital use, flagged patients followed up17,44642.2% of alarms proved correct
Low-dose scans at routine check-ups10,95999.94% specificity; eight scans flagged

The four tests in the table are the main ones. The study’s total of 80,612 also takes in the 35,402 patients used to adjust the system and several smaller groups.

In the main test, on scans from eight hospitals that played no part in training, the system achieved 98.5% specificity, with 90.0% sensitivity for cancer and 52.5% for precancerous lesions.

The early end of the disease is where screening earns its keep, and there the numbers are weaker. In stage 1 cancer, sensitivity was 60.1%. In a separate group of 702 people who had both a scan and an endoscopy, with the system set to be more sensitive at the cost of more false alarms, sensitivities were 65.0% for precancerous lesions and 78.4% for stage I cancer.

The researchers also tested the system as an assistant. Seventeen radiologists interpreted the same 300 scans twice, first alone and then, at least three months later, with the software’s output in front of them. Their average sensitivity rose from 71.9% to 85.7%, while specificity increased from 79.6% to 91.7%. EAGLE alone performed better than every one of them, Ynetnews reported.

One analysis looked backward. The team went back to old scans from 28 patients who were later diagnosed with esophageal cancer. The system flagged 18 of them, and in four cases the scan had been performed at least nine months before the disease was diagnosed.

False alarms in AI screening of CT scans

A specificity of 98.5% sounds close to perfect. In screening it is the number that decides whether a test is usable. Of every 1,000 healthy people, only about 15 were incorrectly flagged as suspicious. Where a cancer is rare, those 15 can outnumber the real cases several times over, and each could mean an endoscopy that turns out to be unnecessary.

The same research group made this point about its earlier pancreatic model. Screening people without symptoms with a single test, that paper said, has been held back by the low prevalence and potential harms of false positives.

The team addressed it by adjusting the system on real-world data. Calibration in a group of 35,402 patients reduced false positives by 72.7% while preserving sensitivity. In the hospital test that followed, 42.2% of alarms proved correct. More than half, in other words, were false.

Low-dose scans of the kind used in lung cancer screening were tested separately. In a group of 1,607 such scans, the system achieved 88.4% sensitivity and 99% specificity.

In the setting closest to true screening the alarms were rare. Among 10,959 people ages 45 to 75 who underwent low-dose CT as part of routine medical examinations, the system flagged only eight scans as suspicious. One of the eight carried a finding the original reading had missed, and eight days later the patient was diagnosed with esophageal cancer.

Low-dose scans matter because programs that use them already exist. In the United States, lung cancer screening by this method is recommended each year in adults aged 50 to 80 years who have a 20 pack-year smoking history and who still smoke or quit within the past 15 years. A pack-year is one pack of cigarettes a day for a year. Smoking raises the risk of both cancers, so the people being scanned for one are at higher risk of the other.

An outside radiologist on AI and opportunistic screening

Arnon Makori, head of imaging at Assuta Medical Centers in Israel, was not involved in the work and discussed it with Ynetnews. Makori welcomed the idea of getting more from scans already taken, and was plain about the early-stage figure. “A sensitivity of 60% is not very high,” Makori said, adding that the tool is still maturing.

Makori also questioned how far the results travel. “This is highly relevant to the East Asian population and less relevant to Israel and the Western population in general,” Makori said.

The reason is the biology. In East Asia, squamous cell carcinoma is more common, while in Western countries adenocarcinoma accounts for a larger share of cases. Squamous cell carcinoma arises in the flat cells lining the esophagus, and adenocarcinoma in gland cells, usually near the stomach. Worldwide the first type dominates: a global estimate for 2020 found that 85% of esophageal cancers were squamous cell carcinomas and 14% were adenocarcinomas. In that estimate the highest rates occurred in Eastern Asia and Southern and Eastern Africa.

Makori’s view of the technology was as an aid. Makori described AI as able to serve “as a co-pilot, navigator or decision-support system for the radiologist interpreting the scan”.

What is not yet known about AI screening for esophageal cancer

Whether it saves lives. The study measured detection. The authors’ own claim is modest: the system has the potential to serve as a scalable tool for early screening. Showing that people screened this way are less likely to die of the disease would take a different and much longer study.

Whether it works outside China. The system was trained entirely on data from China, a country that accounts for roughly half of the esophageal cancer cases in the world. Centers in two other countries took part in the testing. The system was less effective at identifying tumors near the junction of the esophagus and stomach, which is where adenocarcinoma usually arises.

Women. The system performed better in men than in women. The disease is far more common in men, with incidence and mortality rates 2- to 3-fold higher in male than in female populations worldwide. The researchers suggest that this may be why: the training data held fewer women.

Small numbers where it counts. The real-world tests included few cancers. In some of them, follow-up lasted less than two years and not everyone flagged as suspicious completed further testing, so missed cancers and unconfirmed alarms may both be undercounted.

Using it to ration endoscopy. The paper suggests that referring high-risk individuals for endoscopy could improve screening efficiency, and a simulation put the share of endoscopies avoided at about two-thirds. The authors label those analyses exploratory.

Who built it. The system was developed by the research arm of a technology company, which announced it alongside its other cancer-detection models. It is a research tool, and the study does not make it an approved screening test.

Whether any screening test is appropriate depends on a person’s own risk, which is assessed by their doctor.

An AI system detected 90.0% of esophageal cancers in its main test on ordinary chest CT scans while wrongly flagging about 15 in every 1,000 people without the disease, in a study spanning more than 80,000 people, a result that comes from a system trained in China, weaker for the earliest lesions, and not yet shown to reduce deaths.

People also ask

What is opportunistic screening?

Using a scan that was ordered for one reason to look for another disease. A chest CT taken to examine the lungs also captures the esophagus, so software can check it without another test or extra radiation.

How accurate was the system?

In tests on 11,466 patients at eight centers, it detected 90.0% of esophageal cancers and 52.5% of precancerous lesions, with 98.5% specificity. Detection of stage 1 cancer, the earliest stage, was 60.1%.

What does 98.5% specificity mean?

Of every 1,000 people without the disease, about 15 were wrongly flagged. In a hospital test of 17,446 patients, 42.2% of alarms proved correct, so more than half were false.

Does it work for the type of esophageal cancer common in Western countries?

That is uncertain. The system was trained entirely on data from China, where squamous cell carcinoma is the usual type. It was less effective for tumors near the junction with the stomach, where the adenocarcinoma more common in Western countries tends to arise.

Can a chest CT now be used to screen for esophageal cancer?

No. The system is not an approved screening test and does not replace endoscopy. The study measured how well it finds tumors, not whether using it reduces deaths. Whether any test is appropriate for a particular person is assessed by their doctor. This is general information rather than medical advice.

References

  1. Zhou, J., Guo, G., Yao, J., et al. Large-scale esophageal cancer screening through noncontrast computed tomography and artificial intelligence. Nature Medicine, 2026.
  2. Gefen, E. AI spots esophageal cancer on CT scans months before diagnosis. Ynetnews, 2026.
  3. Alibaba Cloud Community. Alibaba DAMO Academy Unveils New AI Models to Advance Cancer Diagnoses. 2026.
  4. Morgan, E., Soerjomataram, I., Rumgay, H., et al. The Global Landscape of Esophageal Squamous Cell Carcinoma and Esophageal Adenocarcinoma Incidence and Mortality in 2020 and Projections to 2040. Gastroenterology, 2022.
  5. Chen, R., Liu, Y., Song, G., et al. Effectiveness of one-time endoscopic screening programme in prevention of upper gastrointestinal cancer in China: a multicentre population-based cohort study. Gut, 2020.
  6. Cao, K., Xia, Y., Yao, J., et al. Large-scale pancreatic cancer detection via non-contrast CT and deep learning. Nature Medicine, 2023.
  7. US Preventive Services Task Force. Screening for Lung Cancer: US Preventive Services Task Force Recommendation Statement. JAMA, 2021.
Search