Skip to main content ->
Ai2

Research - Papers

Explore a selection of our published work on a variety of key research challenges in AI.

Filter papers

Rethinking the Evaluation of Harness Evolution for Agents

Yike WangHuaisheng ZhuZhengyu HuTeng Xiao
2026
arXiv

We revisit the evaluation of automatic harness evolution for LLM agents. Existing harness evolution methods use unit test cases to search for harness configurations and then report final performance… 

Narrative Scaffolding: A Narrative-First Framework for Data-Driven Sensemaking

Oliver HuangMuhammad FatirTian LuoCarolina Nobre
2026
International Conference on Intelligent User Interfaces (IUI)

When exploring data, analysts construct narratives about what the data means by asking questions, generating visualizations, reflecting on patterns, and revising their interpretations as new… 

AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use

Junzhi ChenHarsh TrivediJane PanAshish Sabharwal
2026
ICML

Tool-use agents that address day-to-day digital tasks such as ordering groceries must not only operate applications, but also interact with the user, e.g., to ask clarification questions, prompt for… 

Demystifying Scientific Problem-Solving in LLMs by Probing Knowledge and Reasoning

Alan LiYixin LiuArpan SarkarArman Cohan
2026
ICML

Scientific problem solving poses unique challenges for LLMs, requiring both deep domain knowledge and the ability to apply such knowledge through complex reasoning. While automated scientific… 

DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research

Rulin ShaoAkari AsaiShannon Zejiang ShenPang Wei Koh
2026
ICML 2026

Deep research agents perform multi-step research to produce long-form, well-attributed answers. However, most open deep research agents are trained on easily verifiable short-form QA tasks via… 

Revisiting Padded Transformer Expressivity: Which Architectural Choices Matter and Which Don't

Anej SveteWilliam MerrillRyan CotterellAshish Sabharwal
2026
ICML

Recent work describes what transformers can and cannot compute through connections to boolean circuits, but existing results lack exact characterizations and are sensitive to modeling choices.… 

Why Are Linear RNNs More Parallelizable?

William MerrillHongjian JiangYanhong LiAshish Sabharwal
2026
ICML

The community is increasingly exploring linear RNNs (LRNNs) as language models, motivated by their expressive power and parallelizability. While prior work establishes the expressivity benefits of… 

Language Models as Higher-Order Planning Formalizers

Owen JiangCassie HuangAshish SabharwalLi Zhang
2026
arXiv.org

Recent work provides overwhelming evidence that LLMs, even those trained to scale their reasoning trace, quickly deteriorate at planning as problems become more complex. LLM-as-Formalizers aim to… 

Generating Literature-Driven Scientific Theories at Scale

Peter JansenPeter ClarkDoug DowneyDaniel S. Weld
2026
ACL

Contemporary automated scientific discovery has focused on agents for generating scientific experiments, while systems that perform higher-level scientific activities such as theory building remain… 

Language Models Don't Know What You Want: Evaluating Personalization in Deep Research Needs Real Users

Nishant BalepurMalachi HamadaV. KishoreAakanksha Naik
2026
ACL

Deep Research (DR) systems help researchers cope with ballooning publishing counts. Such tools synthesize scientific papers to answer research queries, but lack understanding of their users. We…