Ai2 Newsletter
September 2026
Top story - AutoDiscovery surfaced an unexpected cancer signal that Providence researchers went on to confirm
AI’s biggest role in science may not be answering questions. It may be helping scientists identify which questions are worth asking.
That’s the idea behind AutoDiscovery, our system for open-ended scientific exploration. A joint Providence-Ai2 team led by Dr. Kelly Paulson of the Paul G. Allen Research Center (PARC), Dr. Sasha Stanton of the Earle A. Chiles Research Institute, and Ai2 Senior Research Scientist Bodhisattwa Majumder put it to work on The Cancer Genome Atlas, one of the world’s most studied cancer datasets. AutoDiscovery generated and tested hypotheses across the data, looking for patterns that challenged expectations.
Watch our video featurette here.
One AutoDiscovery-generated hypothesis involved invasive lobular carcinoma (ILC), a breast cancer subtype affecting roughly 48,000 Americans each year. ILC has long been considered relatively “immune cold,” but AutoDiscovery surfaced more immune activity than expected.
Follow-up research confirmed the signal. The team found the same pattern in a separate patient dataset, then analyzed ILC tumor samples in the lab and found T-cells around the tumors. Together, the evidence suggests ILC warrants broader investigation for immunotherapy.

Now we’re expanding the collaboration. Ai2 and Providence Swedish Cancer Institute are partnering to deploy AutoDiscovery across additional cancer research datasets at PARC, including protected data within Providence’s own cloud environment.

Ultimately, the promise is to expand what scientists are able to discover—using AI to explore questions and patterns at a scale no person could alone, while scientists bring the judgment needed to decide what is meaningful, trustworthy, and worth pursuing.
If we get that relationship right, it could open new directions for research and reveal connections that might otherwise go unnoticed, while keeping scientists in control.
TutorMoments
We introduced a preview of TutorMoments, an eval for testing how well LLMs tutor students–-e.g., when to give more help and when to let students do the reasoning themselves. Built from one-on-one math tutoring sessions, it includes data and tools for others to study and improve AI tutoring.
Building better Thai LLMs with Dolma
Thai researchers used our open Dolma pretraining pipeline to build Mangosteen, a 47-billion-token Thai dataset designed around the language’s specific needs and cultural nuances.
Studying model drug morphology with Olmo
Researchers studying Olmo 3 found that models can sometimes infer a drug’s properties from patterns in its name rather than drug knowledge. Because Olmo’s data is open, the team could trace that behavior back to how frequently individual drugs appeared during training.
BenchMIRT
Our new BenchMIRT framework audits LLM evaluations question by question to reveal which underlying capabilities are actually driving their scores. Across 100 LLMs, 16 benchmarks, and more than 34,000 questions, it found cases where evaluations designed around safety were also strongly shaped by general reasoning—and identified which questions carried the most useful signal.
New: OlmoEarth on LinkedIn
OlmoEarth now has a home on LinkedIn—a behind-the-scenes look at what the team is building, from model releases and platform updates to new capabilities, engineering write-ups, partner deployments in the field, and where our state-of-the-art Earth observation models are headed next.