Skip to main content ->
Ai2

Research - Papers

Explore a selection of our published work on a variety of key research challenges in AI.

Filter papers

AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite

Jonathan BraggMike D'ArcyNishant BalepurDaniel S. Weld
2026
ICLR

AI agents hold great real-world promise, with the potential to revolutionize scientific productivity by automating literature reviews, replicating experiments, analyzing data, and even proposing new… 

Improving Attributed Long-form Question Answering with Intent Awareness

Xinran ZhaoAakanksha NaikJay DeYoungV. Kishore
2026
arXiv.org

Large language models (LLMs) are increasingly being used to generate comprehensive, knowledge-intensive reports. However, while these models are trained on diverse academic papers and reports, they… 

On the Reasoning Abilities of Masked Diffusion Language Models

Anej SveteAshish Sabharwal
2026
ICLR

Masked diffusion models (MDMs) for text offer a compelling alternative to traditional autoregressive language models. Parallel generation makes them efficient, but their computational capabilities… 

Probabilistic Programs of Thought

Poorva GargRenato Lui GehDaniel Mingyi IsraelGuy Van den Broeck
2026
arXiv

LMs are widely used for code generation and mathematical reasoning tasks where they are required to generate structured output. They either need to reason about code, generate code for a given… 

Cocoa: Co-Planning and Co-Execution with AI Agents

K. FengKevin PuMatt LatzkeJoseph Chee Chang
2026
Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems

As AI agents take on increasingly long-running tasks involving sophisticated planning and execution, there is a corresponding need for novel interaction designs that enable deeper human-agent… 

Perspectra: Choosing Your Experts Enhances Critical Thinking in Multi-Agent Research Ideation

Yiren LiuViraj ShahSangho SuhYun Huang
2026
Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems

Early-stage interdisciplinary research ideation is often challenged by limited expert access, uncertainty about what to ask, and the cognitive burden of synthesizing unfamiliar domain perspectives.… 

Omakase: proactive assistance with actionable suggestions for evolving scientific research projects

Pao SiangliulueJonathan BraggDoug DowneyDaniel S. Weld
2026
arXiv

As AI agents become increasingly capable of complex knowledge tasks, the lack of context limits their capability to proactively reason about a user's latent needs throughout a long evolving project.… 

FloeNet: A mass-conserving global sea ice emulator that generalizes across climates

William GregoryM. BushukJames P. C. DuncanL. Zanna
2026
arXiv

We introduce FloeNet, a machine-learning emulator trained on the Geophysical Fluid Dynamics Laboratory global sea ice model, SIS2. FloeNet is a mass-conserving model, emulating 6-hour mass and area… 

PreScience: A Dataset and Benchmark for Scientific Forecasting

Anirudh AjithAmanpreet SinghJay DeYoungDoug Downey
2026
arXiv

Can AI systems trained on the existing scientific record forecast the advances that will follow? We introduce PreScience, a dataset and benchmark for scientific forecasting built around 98K recent… 

Examining Fast Radiative Feedbacks Using Machine-Learning Weather Emulators

Ankur MaheshWilliam D. CollinsTravis A. O'BrienDa Yang
2026
arXiv

The response of the climate system to increased greenhouse gases and other radiative perturbations is governed by a combination of fast and slow feedbacks. Slow feedbacks are typically activated in…