Immunotherapy - a cancer treatment that boosts and uses the power of the patient’s immune system against cancer cells - has transformed lung cancer treatment, achieving long-term benefit in 20% to 30% of patients. However, most patients experience resistance to treatment, and physicians still cannot reliably predict which patients will benefit from immunotherapy. Today, treatment decisions rely heavily on the expression of PD-L1 - a biomarker with well-known limitations.
“We need smarter tools,” said thoracic oncologist Marina Garassino, MD, Professor of Medicine at UChicago Medicine and senior author of the study.
Better predictive biomarkers could identify which patients are likely to respond to immunotherapy at diagnosis, avoiding unnecessary toxicity and unnecessary cost. Improving the ability to make these predictions can help physicians better tailor treatments for individual patients. The I3LUNG project set out to develop and validate AI tools to support immunotherapy decisions in advanced NSCLC.
For the study, the international research team enrolled 2,396 patients with advanced NSCLC treated with immunotherapy across six centers in Italy, Germany, Greece, Israel, Spain and the United States. For the first phase of the project, the team integrated multiple types of data, including clinical, imaging, pathology, and genomic data, from each patient into a database. They then built and tested two families of AI models trained to predict treatment response and survival using all the collected data.
The researchers found that their AI models consistently outperformed all standard clinical biomarkers. Area Under the Curve, or AUC, is a machine learning metric that quantifies how good an AI model is at classifying information correctly, or in this case, predicting survival outcomes accurately. Scores between 0.8 and 0.9 are considered “excellent.” In the I3LUNG study, the AI model that used clinical and blood data achieved an AUC score of 0.77, while the model that incorporated clinical and blood data along with imaging and digital pathology achieved an AU score of 0.88.
Part of the study tested what happens when humans collaborate with the AI tools. Twenty physicians - 10 lung cancer experts and 10 from other specialties - reviewed 100 real patient cases, first without and then with AI support. Access to the AI tool improved sensitivity for identifying responders from an AUC of 0.72 to 0.87. Physicians who were not lung cancer experts showed the greatest improvements, a finding with direct relevance to community oncology settings where thoracic expertise may be limited. Inter-physician agreement rose from slight to moderate, suggesting the tool also promotes more consistent clinical reasoning across different levels of experience.
“This alignment between machine and clinical logic is essential for building trust in AI-assisted decision-making,” Garassino said.
The current study reported the project’s retrospective phase. I3LUNG is now prospectively enrolling more than 2,000 patients across the same six international centers, allowing the researchers to focus on treatment optimization. They said this dimension is important because AI model evaluation should address not only model performance, but also its usability in the clinic.
“I3LUNG establishes a new benchmark for AI in thoracic oncology. Decision support tools built even from routinely available clinical data can outperform the biomarkers we rely on today,” Garassino said. “For patients, this means fewer missed opportunities for treatment benefit. For community physicians, it means access to expert-level guidance at the point of care. For the field, it provides a rigorous, fair, and explainable framework - validated across diverse healthcare systems and populations - that can serve as a global platform for the next generation of precision immunotherapy.”
The study, “I3LUNG: Clinical Validation of a Multimodal AI Tool to Support Immunotherapy Decisions in NSCLC,” was supported by funding from the European Union’s Horizon 2020 research and innovation program.
Prelaj A, Miskovic V, Sacco M, Ferrarin A, Licciardello CM, Provenzano L, Favali M, Lerma L, Zec A, Spagnoletti A, Ganzinelli M, Lorenzini D, Guirges B, Invernizzi L, Silvestri C, Mazzeo L, Prina MM, Corrao G, Ruggirello M, Dumitrascu AD, Di Mauro RM, Monzani D, Pravettoni G, Zanitti M, Macocchi D, Marino MB, Cavalli C, Romanò R, Giani C, Armato SG 3rd, Esposito A, Bestvina CM, Spector M, Naama BR, Basheer R, Hafzadi AL, Roisman L, Watermann I, Szewczyk M, Olchers T, Richter H, Blanke-Roeser C, Siniscalchi C, Di Lello A, Arangoa T, Bartolomeo V, Spathas N, Sarris E, Fountzilas E, Roca AA, Caro-Consuegra R, Iranzo P, Fernández-Pinto M, Rodríguez-Morató J, Agnelli L, Occhipinti M, Brambilla M, Beninato T, Proto C, Kosta S, Di Palma MP, Rulli E, Steurer S, Simon R, Willis M, Pruneri G, De Braud F, Restelli M, Felip E, Peled N, Pearson AT, Linardou H, Reck M, Russo GL, Trovò F, Pedrocchi ALG, Garassino MC; I3LUNG Consortium.
Clinical usability of an explainable AI decision support tool and evaluation of multimodal models in NSCLC.
Nat Med. 2026 Sep;32(9):3235-3247. doi: 10.1038/s41591-026-04488-2