Skip to main content ->
Ai2

Research - Papers

Explore a selection of our published work on a variety of key research challenges in AI.

Filter papers

LitPivot: Developing Well-Situated Research Ideas Through Dynamic Contextualization and Critique within the Literature Landscape

Hita KambhamettuBhavana Dalvi MishraAndrew HeadPao Siangliulue
2026
UIST 2026

Developing a novel research idea is hard. It must be distinct enough from prior work to claim a contribution while also building on it. This requires iteratively reviewing literature and refining an… 

Olmo Hybrid: From Theory to Practice and Back

William MerrillYanhong LiTyler RomeroAshish Sabharwal
2026
COLM

Recent work has demonstrated the potential of non-transformer language models, especially linear recurrent neural networks (RNNs) and hybrid models that mix recurrence and attention. Yet there is no… 

From Token Probabilities to Semantic Constraints: Towards Declarative Probabilistic Evaluation of Language Models

Kyle RichardsonCullen AndersonPranav BalakrishnanMarisa Hudspeth
2026
EMNLP

While Large Language Models have improved rapidly, many fundamental questions remain about how to evaluate the knowledge and reasoning abilities they acquire, and how such evaluations relate to the… 

Process-Oriented Evaluation of AI-Assisted Scientific Writing

Patrick Queiroz Da SilvaSanchaita HazraDoeun LeeBodhisattwa Prasad Majumder
2026
Conference on Language Modeling

Bad writing hinders the publication of science. The role of artificial intelligence (AI) in generating and editing scientific texts remains unsettled. Abstracts serve as the critical gateway to… 

ArtifactLinker: Linking Scientific Artifacts for Automatic State-of-the-Art Discovery

Haofei YuJiaxuan YouPeter ClarkKyle Richardson
2026
COLM 2026

Scientific artifacts such as models and datasets are foundations for research. With the rapid growth of platforms like HuggingFace, researchers now have access to a large number of artifacts. Yet, a… 

CoTs as Tractable Probabilistic Programs

Kyle RichardsonYu FengPoorva GargDan Roth
2026
9th Workshop on Tractable Probabilistic Modeling (TPM@UAI)

Chain-of-thought (CoT) traces are used across language model prompting, training, test-time inference, and interpretability, yet they are often modeled in task-specific ways. We propose treating CoT… 

AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use

Junzhi ChenHarsh TrivediJane PanAshish Sabharwal
2026
ICML

Tool-use agents that address day-to-day digital tasks such as ordering groceries must not only operate applications, but also interact with the user, e.g., to ask clarification questions, prompt for… 

Revisiting Padded Transformer Expressivity: Which Architectural Choices Matter and Which Don't

Anej SveteWilliam MerrillRyan CotterellAshish Sabharwal
2026
ICML

Recent work describes what transformers can and cannot compute through connections to boolean circuits, but existing results lack exact characterizations and are sensitive to modeling choices.… 

Why Are Linear RNNs More Parallelizable?

William MerrillHongjian JiangYanhong LiAshish Sabharwal
2026
ICML

The community is increasingly exploring linear RNNs (LRNNs) as language models, motivated by their expressive power and parallelizability. While prior work establishes the expressivity benefits of… 

Language Models as Higher-Order Planning Formalizers

Owen JiangCassie HuangAshish SabharwalLi Zhang
2026
arXiv.org

Recent work provides overwhelming evidence that LLMs, even those trained to scale their reasoning trace, quickly deteriorate at planning as problems become more complex. LLM-as-Formalizers aim to…