Skip to main content ->
Ai2

Research - Papers

Explore a selection of our published work on a variety of key research challenges in AI.

Filter papers

MolmoAct2: Action Reasoning Models for Real-world Deployment

Haoquan FangJiafei DuanD. ClayRanjay Krishna
2026
CoRL

Vision-Language-Action (VLA) models aim to provide a single generalist controller for robots, but today's systems fall short on the criteria that matter for real-world deployment. Frontier models… 

MolmoB0T: Large-Scale Simulation Enables Zero-Shot Manipulation

Abhay DeshpandeM. GuruRose HendrixRanjay Krishna
2026
CoRL, ICRA • SDRL Workshop, ICRA • VLA Pipeline Workshop, ICRA • Beyond Teleoperation Workshop

A prevailing view in robot learning is that simulation alone is not enough; effective sim-to-real transfer is widely believed to require at least some real-world data collection or task-specific… 

HARPA: A Testability-Driven, Literature-Grounded Framework for Research Ideation

Rosni VasuPeter JansenPao SiangliulueBhavana Dalvi
2026
AACL-IJCNLP 2026

While there has been a surge of interest in automated scientific discovery (ASD), especially with the emergence of LLMs, it remains challenging for tools to generate hypotheses that are both… 

LitPivot: Developing Well-Situated Research Ideas Through Dynamic Contextualization and Critique within the Literature Landscape

Hita KambhamettuBhavana Dalvi MishraAndrew HeadPao Siangliulue
2026
UIST 2026

Developing a novel research idea is hard. It must be distinct enough from prior work to claim a contribution while also building on it. This requires iteratively reviewing literature and refining an… 

Context-Aware RL for Agentic and Multimodal LLMs

Peiyang XuBangzheng LiSijia LiuXingyu Fu
2026
COLM • Efficient Reasoning (spotlight)

Large language models (LLMs) often fail when answering requires identifying a small but decisive piece of evidence within a long or complex context, such as a single line in a tool trace or a subtle… 

Cracks in the Foundation: Seemingly Minor Architectural Choices Impact Long Context Extension

Amanda BertschLuca SoldainiMatthew R. GormleyD. Groeneveld
2026
COLM

One might imagine that architectural variations within the dense transformer paradigm have a limited effect on accuracy. However, we demonstrate that this is not the case in the long context… 

Olmo Hybrid: From Theory to Practice and Back

William MerrillYanhong LiTyler RomeroAshish Sabharwal
2026
COLM

Recent work has demonstrated the potential of non-transformer language models, especially linear recurrent neural networks (RNNs) and hybrid models that mix recurrence and attention. Yet there is no… 

Measuring AI Scientists: From Exams to Discovery

Yuanqi DuSteven DillmannJon LaurentChenru Duan
2026
arXiv

Large language models and agentic systems are increasingly embedded across the scientific work-flow, from literature synthesis and hypothesis generation to code execution, data analysis and writing.… 

Process-Oriented Evaluation of AI-Assisted Scientific Writing

Patrick Queiroz Da SilvaSanchaita HazraDoeun LeeBodhisattwa Prasad Majumder
2026
Conference on Language Modeling

Bad writing hinders the publication of science. The role of artificial intelligence (AI) in generating and editing scientific texts remains unsettled. Abstracts serve as the critical gateway to… 

Stochastic Emulation of a Fully Coupled Preindustrial E3SMv3 Simulation

E. WuJames P. C. DuncanTroy ArcomanoPeter M. Caldwell
2026
arXiv

We present a stochastic coupled emulator of E3SM version 3, built on the SamudrACE framework, which couples an atmosphere emulator (ACE2) with a full-depth ocean emulator (Samudra). We replace the… 

1-10Next