Skip to main content ->
Ai2

Teaching future scientists to interrogate AI tools for scientific discovery

September 14, 2026

Ai2


AutoDiscovery, our AI agent for scientific research, analyzes datasets, proposes hypotheses, runs experiments to test them, and ranks the results by Bayesian surprise—the gap between what the model predicted and what the data showed. In a University of Washington classroom this spring, it surfaced a possible flaw in how battery aging is measured, an unexpected pattern in silicon polymers, and a recurring feature in proteins that respond to light.

But a surprising result is not necessarily an important one—or one that will withstand further testing. Students still had to scrutinize the evidence, compare the findings with published research, and decide whether each result reflected a genuine discovery, a coincidence, or a problem with the data.

Luna Yue Huang, an associate teaching professor of materials science and engineering at the UW, invited students to explore this way of working with AI through the GenAI for Science: Ai2–UW Materials Challenge. Twenty-five teams proposed a dataset and research question; ten were selected to spend several weeks working with AutoDiscovery.

A UW instructional team and Ai2 researchers reviewed the proposals based on the richness and scientific distinctiveness of the data, their potential to produce meaningful findings, and the overall diversity of topics represented. Ten projects were selected, with students from other proposals joining those teams. In some cases, instructors also helped narrow broad datasets to a more manageable scope aligned with the students’ research interests.

Students tested AutoDiscovery’s hypotheses against the available evidence, identified gaps in its reasoning, and assessed whether it meaningfully reduced the amount of manual data analysis required. AutoDiscovery's performance varied considerably—in some cases, it surfaced leads the students might not otherwise have explored.  

Here are four representative projects from the class:

Project 1: Lithium battery aging

A group studying lithium-ion battery aging surfaced a counterintuitive lead: that the periodic diagnostic tests used to check a cell's health may themselves speed its degradation. This rested on a correlation in the data rather than proof that the tests caused the extra wear, so confirming it would take a controlled experiment—a clear and logical next step for follow-on work.

Project 2: Estimating an atom’s charge

Another group, drawing on six million polymer simulations, found an apparent exception to a decades-old rule: that the two standard ways of estimating an atom's charge disagree most on electron-hungry atoms like oxygen and nitrogen. They traced it to the original paper documenting the pattern and found no published explanation for why polymers would diverge from it—a genuine open question.

Project 3: Building light-sensitive proteins

A third group was building proteins that switch on only under light, which means inserting a light-sensitive segment at exactly the right spot in a protein—slow work to test one design at a time in the lab. So the group used AutoDiscovery to find characteristics that separated the insertions that worked from those that didn't, then checked the tool's best hypothesis against their own lab results. The working insertions, AutoDiscovery found, anchored the light-sensitive segment to the protein at a few strong points rather than many weak ones—a pattern that held across every design the group had tested in the real world and matched a long-established principle of protein binding.

Project 4: Examining nuclear reactor configurations

Other projects showed how much the output depended on the data. Working with 24,000 simulated nuclear-reactor configurations, one student group got almost nothing usable out of AutoDiscovery—roughly half its hypotheses were illogical, and none were worth pursuing. The group traced the problem to how AutoDiscovery works; the system gauges surprise against real-world measurements, so the synthetic numbers threw it off.

On a dataset of real reactor-physics measurements – readings from experiments rather than a simulation – AutoDiscovery did far better, turning up about 15 leads potentially worth chasing. Triaging the most promising was where the students came in—weighing each for physical plausibility and logical soundness, checking the strongest against the literature, and deciding which ones warranted a closer look.

The takeaways

Across the projects, the students reached the same conclusion: AutoDiscovery was most useful for surfacing ideas they might not have thought to test, but it didn’t take the researcher out of the loop—if anything, it demanded more of them. Reading its evidence and working out how it reached each result needed the kind of judgment that comes only from experience and scientific expertise.

For Huang, the exercise pointed back at her own teaching—what it takes to prepare students to use a tool like AutoDiscovery well, and to teach the judgment the tool can't replace.

“What made this challenge especially meaningful was that it exposed students to a very different way of working with AI,” says Huang. “Rather than asking AI to complete a well-defined task, they learned to use it as a partner in scientific exploration, with AI generating hypotheses, and students questioning unexpected results, and deciding which ideas deserved further investigation. That shift naturally led students to ask an even more important question: What is the role of the researcher when AI can generate ideas so quickly?” 

“Our goal was to help students recognize that the human contribution becomes even more valuable in this new paradigm,” continues Huang. “Critical thinking, systematic experimental design, scientific reasoning, and the ability to justify conclusions remain uniquely human responsibilities. Experiences like this not only prepare students for a future in which AI becomes an integral part of scientific discovery, but also reinforce the value of scientific education by demonstrating that foundational knowledge, domain expertise, and sound judgment become even more important in the age of AI.”

Huang was also struck by the amount and quality of work the students completed during the challenge.

“Within just ten days, the students transformed raw datasets into carefully investigated scientific hypotheses, supported by literature reviews, quantitative analysis, experimental validation plans, and thoughtfully reasoned reports,” says Huang. “AI undoubtedly accelerated many parts of this process, but the quality of the final work reflected the students’ curiosity, persistence, and scientific judgment.”

“What was especially encouraging was seeing how eager they were to learn not just how to use AI, but how to conduct rigorous research alongside it. For me, that was a powerful reminder that the next generation of scientists is ready to embrace AI as a tool for discovery while recognizing that meaningful scientific progress still depends on critical thinking, evidence-based reasoning, and deep domain expertise.”

We're grateful to the students and faculty in the UW's Department of Materials Science & Engineering for the rigor they brought to testing AutoDiscovery. Their evaluations help us improve it—and show us how tools like it are changing scientific practice, which for us is part of building them responsibly.

Try AutoDiscovery in Asta.

Join us

At Ai2 we’re building the future of transparent, open-source AI — built in the open to empower scientific progress and fundamental understanding of this world changing technology. We’re not here to make profits, we’re here to make sure benefits of AI are shared widely and for the benefit of humanity. If this appeals to you, please take a look at our open roles.

Subscribe to receive monthly updates about the latest Ai2 news.