Immunotherapy represents one of the most significant developments in modern cancer medicine. Yet predicting whether a patient will actually respond remains challenging: In many common tumor types, only a fraction of unselected patient groups benefit from treatment. The rest endure severe side effects without therapeutic benefit. A team from Harvard Medical School and Roche presented on July 3, 2026, in Nature Medicine an AI model that could finally answer this critical question reliably.
The immunotherapy problem: who responds and who doesn't
Immune checkpoint inhibitors (ICI) activate the immune system to recognize and fight cancer cells. The principle works spectacularly well in some patients. In others, it shows little response. Previous tests like PD-L1 tissue staining or tumor mutational burden provide only rough guidance about who truly responds.
This has concrete consequences: ICI therapies cost tens of thousands of euros per treatment cycle and can cause months of severe side effects, including colitis, pneumonitis, and autoimmune organ inflammation. Patients who don't respond lose valuable time and health that could be used for more effective alternatives. With over 500,000 new cancer diagnoses annually, immunotherapy is increasingly the standard treatment for a growing proportion.
How COMPASS opens the black box
The model is called COMPASS and was developed by a team led by Marinka Zitnik from Harvard Medical School and Daniel Marbach from Roche Innovation Center Basel. The crucial difference from previous AI approaches in oncology lies in its architecture: a concept bottleneck transformer.
Conventional AI models process input data—here, activity patterns of roughly 16,000 genes from tumor tissue—and directly deliver an output: respond or don't respond. The computation in between remains hidden. COMPASS forces a biologically interpretable detour: the model must formulate its prediction through 44 biologically defined concepts, including immune cell states in the tumor microenvironment, cancer-immune signaling pathways, and structural properties of the tumor microenvironment. Clinicians see not only whether a patient is likely to respond, but also which biological mechanisms led to this prediction.
This is no academic detail. A doctor who understands the reasoning of an AI model can evaluate it, question it, and explain it to the patient. Previous AI systems in medicine were criticized as black boxes that deliver results but provide no interpretable decision rationale. COMPASS represents a concrete attempt to address this critique.
8.5 percent better than 22 comparison methods
COMPASS was first trained on 10,184 tumors from the Cancer Genome Atlas, covering 33 cancer types. The team then refined the model on 16 clinical ICI cohorts representing seven cancer types and six different therapy regimens, including anti-PD-1, anti-PD-L1, and anti-CTLA4 treatments.
In direct comparison with 22 existing prognostic methods, COMPASS achieved an improvement of 8.5 percent in AUROC (Area Under the Receiver Operating Characteristic Curve) over the previously best method and 15.7 percent in Area Under Precision-Recall Curve, according to the Nature Medicine publication. Particularly relevant: COMPASS also generalized to cancer types and therapy regimens not represented in training. This transferability to unseen scenarios rarely succeeds with previous models.
What the study doesn't establish: whether COMPASS performs equally well in prospective clinical trials. All comparisons are based on retrospective cohort data. The jump from historical training data to real clinical decision support requires separate validation in new patient populations.
Validation at German NCT centers: first results in 12 to 24 months
Zitnik's lab has published the code and a web implementation on GitHub, allowing external research groups to test COMPASS on their own patient data. For clinical use, the model requires RNA sequencing of tumor tissue to provide gene expression profiles. In Germany, the NCT centers (National Center for Tumor Diseases) in Heidelberg, Dresden, Berlin, and Tübingen, as well as the Comprehensive Cancer Centers at university hospitals, have the necessary infrastructure. In primary care settings, RNA sequencing is not yet routine.
Two caveats remain. First, note the collaboration with Roche: the pharmaceutical company has a commercial interest in optimizing its immunotherapies. External validations by groups without Roche ties are critical before COMPASS should be deployed in patient care. Second, FDA requirements for prospective data apply: retrospective cohort data, as COMPASS relies on, does not meet FDA guidance for AI-based medical devices for clinical approval. The researchers themselves name prospective studies as the next step, where COMPASS predictions actually influence treatment decisions and outcomes are measured. First results are realistically expected within 12 to 24 months.
