Skip to main content ->
Ai2

Research - Papers

Explore a selection of our published work on a variety of key research challenges in AI.

Filter papers

HARPA: A Testability-Driven, Literature-Grounded Framework for Research Ideation

Rosni VasuPeter JansenPao SiangliulueBhavana Dalvi
2026
AACL-IJCNLP 2026

While there has been a surge of interest in automated scientific discovery (ASD), especially with the emergence of LLMs, it remains challenging for tools to generate hypotheses that are both… 

Querying Multimodal Scientific Papers with AI: Practices and Preferences Across Blind, Low-Vision, and Sighted Scientists

Arnavi Chheda-KotharyLucy Lu WangJoseph Chee ChangJonathan Bragg
2026
ASSETS

Visual diagrams, figures, and tables are central to scientific papers, and convey information beyond what is captured in text. While blind or low-vision (BLV) scientists have traditionally relied on… 

Demystifying Scientific Problem-Solving in LLMs by Probing Knowledge and Reasoning

Alan LiYixin LiuArpan SarkarArman Cohan
2026
ICML

Scientific problem solving poses unique challenges for LLMs, requiring both deep domain knowledge and the ability to apply such knowledge through complex reasoning. While automated scientific… 

Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users

Nishant BalepurMalachi HamadaV. KishoreAakanksha Naik
2026
ACL

Deep Research (DR) systems help researchers cope with ballooning publishing counts. Such tools synthesize scientific papers to answer research queries, but lack understanding of their users. We… 

Scientific reasoning does not reliably translate into scientific forecasting in frontier AI

Sean WuPan LuYupeng ChenJunchi Yu
2026
arXiv

AI systems are increasingly used to support forward-looking scientific judgment, but it remains unclear whether they can form reliable expectations about future scientific advances. Here we show… 

Improving Attributed Long-form Question Answering with Intent Awareness

Xinran ZhaoAakanksha NaikJay DeYoungV. Kishore
2026
ICLR

Large language models (LLMs) are increasingly being used to generate comprehensive, knowledge-intensive reports. However, while these models are trained on diverse academic papers and reports, they… 

DRACULA: Hunting for the Actions Users Want Deep Research Agents to Execute

Nishant BalepurMalachi HamadaVarsha KishoreAakanksha Naik
2026
arXiv

Scientific Deep Research (DR) agents answer user queries by synthesizing research papers into multi-section reports. User feedback can improve their utility, but existing protocols only score the… 

AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite

Jonathan BraggMike D'ArcyNishant BalepurDaniel S. Weld
2026
ICLR

AI agents hold great real-world promise, with the potential to revolutionize scientific productivity by automating literature reviews, replicating experiments, analyzing data, and even proposing new… 

Cocoa: Co-Planning and Co-Execution with AI Agents

K. FengKevin PuMatt LatzkeJoseph Chee Chang
2026
Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems

As AI agents take on increasingly long-running tasks involving sophisticated planning and execution, there is a corresponding need for novel interaction designs that enable deeper human-agent… 

Understanding Usage and Engagement in AI-Powered Scientific Research Tools: The Asta Interaction Dataset

Dany HaddadDan BareketJ. ChangDoug Downey
2026
arXiv.org

AI-powered scientific research tools are rapidly being integrated into research workflows, yet the field lacks a clear lens into how researchers use these systems in real-world settings. We present… 

1-10Next