AI can already assist with a surprising amount of scientific work: search literature, analyze datasets, write code, and generate and test hypotheses. The harder problem is building systems scientists can steer as a research project unfolds, especially when new evidence changes which hypotheses to pursue, experiments to run, or questions to investigate next.
That challenge was the focus of an August 27 event at Ai2 marking an expansion of our work with the Paul G. Allen Research Center at Providence Swedish Cancer Institute. We’ve written separately about the cancer research that grew out of that collaboration. At the event, that work served as a starting point for a broader look at what today’s scientific AI still lacks.
The conversation included:
- Kelly Paulson, MD, PhD, Lead, Center for Immuno-Oncology, Paul G. Allen Research Center & Medical Director, Melanoma and Cutaneous Oncology, Providence Swedish Cancer Institute
- Bodhisattwa Prasad Majumder, Senior Research Scientist, Ai2
- Sasha Stanton, MD, PhD, Medical Breast Oncologist & Lead, Cancer Immunoprevention Laboratory, Earle A. Chiles Research Institute, Providence Cancer Institute
- Kyle Travaglini, Assistant Investigator, Allen Institute for Brain Health Accelerator
- Abraham Flaxman, PhD, Professor of Health Metrics Sciences and Global Health, University of Washington & Institute for Health Metrics and Evaluation
- Stephen Salerno, PhD, Assistant Professor, WashU Bursky School of Public Health
- Hoifung Poon, PhD, Chief AI Officer, Recursion
Across the presentations and panel, five challenges kept resurfacing.
1. Leaving room for the human element
Majumder described scientific taste as the judgment that helps a researcher distinguish an interesting result from one that’s trivial, implausible, already understood, or unlikely to lead anywhere. Today’s AI-for-science systems don’t reliably make those distinctions on their own.
The goal isn’t necessarily to transfer that judgment wholesale to an AI system. The challenge is giving researchers better ways to bring their expertise, priorities, and changing understanding of a problem into the system so AI can assist without deciding on its own which directions are worth pursuing.
Ai2’s work with Providence researchers made that distinction clear—AutoDiscovery could surface hypotheses that looked surprising statistically but made little biological or clinical sense without expert context. Once researchers brought their domain knowledge to the process, the system became much more useful.
2. Keeping AI steerable
Scientific work rarely follows a fixed plan. An experiment produces an unexpected result, a new paper changes what researchers know, or a scientist decides to add a dataset, swap in a different tool, or redirect an AI agent toward a finding it hadn’t prioritized.
Majumder argued that current agents remain difficult to steer across long-running investigations. Scientists may need to revise an agent’s instructions, context, or tools as new experimental results and findings from the field come in—all without starting over or retraining the underlying model each time.
For scientific agents, steering their behavior and keeping their knowledge current are closely linked. A useful agent can’t simply follow an initial brief; it has to remain adaptable as the research evolves.
3. Deciding what to delegate
Poon drew a distinction between two ways AI can help scientists: productivity gains and creativity gains. Productivity gains come from taking on work people already know how to do but that’s tedious or time-consuming, such as searching records, structuring information, or synthesizing literature. Those tasks are relatively easy to describe and, importantly, to verify.
Creativity gains are harder to evaluate. Asking AI to surface a mechanism scientists haven’t considered or suggest an unexpected experiment can be scientifically valuable, but those outputs can’t usually be checked as simply as a well-defined task. They often require further analysis and replication to determine whether the idea holds up.
Flaxman offered an example from his work as an editor at the Journal of Privacy and Confidentiality. A researcher used an AI system to test algorithms from his own previously published papers, and the system flagged an error in one of them. After investigating the critique, the researcher concluded the AI was right and asked the journal to retract the paper.
The value wasn’t in accepting the AI’s judgment at face value, but in surfacing something worth interrogating. As Paulson put it: “It’s research, not search. You must discover, and you have to verify.”
4. Avoiding the amplification of bad science
Faster analysis doesn’t fix a poorly designed study or bad data. Salerno described AI in that sense as an amplifier rather than an equalizer—it can make strong experimental design, careful data collection, and sound causal reasoning more powerful, but it can also magnify weak assumptions and flawed methodology.
That makes foundational questions about the research process even more important. Where did the data come from? Why was it collected? Who was included or excluded? Does an apparent relationship make causal or scientific sense?
As AI lets researchers analyze more data and test more ideas, weak assumptions and methodological flaws become more consequential too.
5. Tightening the loop between AI and experiments
One of the more ambitious possibilities discussed was a tighter connection between AI and the lab, rather than an AI scientist working independently.
Travaglini described neuroscience projects involving hundreds of cell types and thousands of changing genes—too many relationships for any researcher to trace through the literature alone. Agents could help synthesize that evidence and prioritize promising hypotheses, and eventually interact directly with laboratory instruments so the results of one experiment help shape the next.
Poon outlined a complementary path: richer computational models of biology. He described “virtual tissue” as a potential bridge between models of individual cells and the farther-reaching goal of a “virtual patient” that could help forecast disease progression or response to treatment.
Together, those ideas point toward scientific AI that participates more fully in the research process—helping researchers decide what to investigate, incorporate new evidence, and connect analysis more closely with laboratory work. Getting there will require not only more capable models, but more adaptable systems, better ways to incorporate new knowledge, and enough flexibility for researchers to steer and verify their work as projects evolve.
Join us
At Ai2 we’re building the future of transparent, open-source AI — built in the open to empower scientific progress and fundamental understanding of this world changing technology. We’re not here to make profits, we’re here to make sure benefits of AI are shared widely and for the benefit of humanity. If this appeals to you, please take a look at our open roles.