Skip to main content ->
Ai2

Research - Papers

Explore a selection of our published work on a variety of key research challenges in AI.

Filter papers

From Token Probabilities to Semantic Constraints: Towards Declarative Probabilistic Evaluation of Language Models

Kyle RichardsonCullen AndersonPranav BalakrishnanMarisa Hudspeth
2026
EMNLP

While Large Language Models have improved rapidly, many fundamental questions remain about how to evaluate the knowledge and reasoning abilities they acquire, and how such evaluations relate to the… 

MolmoWeb: Open Visual Web Agent and Open Data for the Open Web

Tanmay GuptaPiper WoltersZixian MaRanjay Krishna
2026
ECCV 2026

Web agents--autonomous systems that navigate and execute tasks on the web on behalf of users--have the potential to transform how people interact with the digital world. However, the most capable… 

XYBench: Can LLMs Respond Pragmatically to Queries with Misconceptions?

Akhila YerukolaJena D. HwangMing-Qian ZhengM. Sap
2026
EMNLP

When non-expert users ask LLMs for assistance, their queries can often have misconceptions (e.g.,"How do I parse XML with regex?"). In such cases, often referred to as the XY-problem, LLMs must… 

Measuring AI Scientists: From Exams to Discovery

Yuanqi DuSteven DillmannJon LaurentChenru Duan
2026
arXiv

Large language models and agentic systems are increasingly embedded across the scientific work-flow, from literature synthesis and hypothesis generation to code execution, data analysis and writing.… 

Process-Oriented Evaluation of AI-Assisted Scientific Writing

Patrick Queiroz Da SilvaSanchaita HazraDoeun LeeBodhisattwa Prasad Majumder
2026
Conference on Language Modeling

Bad writing hinders the publication of science. The role of artificial intelligence (AI) in generating and editing scientific texts remains unsettled. Abstracts serve as the critical gateway to… 

Stochastic Emulation of a Fully Coupled Preindustrial E3SMv3 Simulation

E. WuJames P. C. DuncanTroy ArcomanoPeter M. Caldwell
2026
arXiv

We present a stochastic coupled emulator of E3SM version 3, built on the SamudrACE framework, which couples an atmosphere emulator (ACE2) with a full-depth ocean emulator (Samudra). We replace the… 

ArtifactLinker: Linking Scientific Artifacts for Automatic State-of-the-Art Discovery

Haofei YuJiaxuan YouPeter ClarkKyle Richardson
2026
COLM 2026

Scientific artifacts such as models and datasets are foundations for research. With the rapid growth of platforms like HuggingFace, researchers now have access to a large number of artifacts. Yet, a… 

CoTs as Tractable Probabilistic Programs

Kyle RichardsonYu FengPoorva GargDan Roth
2026
9th Workshop on Tractable Probabilistic Modeling (TPM@UAI)

Chain-of-thought (CoT) traces are used across language model prompting, training, test-time inference, and interpretability, yet they are often modeled in task-specific ways. We propose treating CoT… 

Querying Multimodal Scientific Papers with AI: Practices and Preferences Across Blind, Low-Vision, and Sighted Scientists

Arnavi Chheda-KotharyLucy Lu WangJoseph Chee ChangJonathan Bragg
2026
ASSETS

Visual diagrams, figures, and tables are central to scientific papers, and convey information beyond what is captured in text. While blind or low-vision (BLV) scientists have traditionally relied on… 

Rethinking the Evaluation of Harness Evolution for Agents

Yike WangHuaisheng ZhuZhengyu HuTeng Xiao
2026
arXiv

We revisit the evaluation of automatic harness evolution for LLM agents. Existing harness evolution methods use unit test cases to search for harness configurations and then report final performance…