Research - Papers
Explore a selection of our published work on a variety of key research challenges in AI.
From Token Probabilities to Semantic Constraints: Towards Declarative Probabilistic Evaluation of Language Models
While Large Language Models have improved rapidly, many fundamental questions remain about how to evaluate the knowledge and reasoning abilities they acquire, and how such evaluations relate to the…
MolmoWeb: Open Visual Web Agent and Open Data for the Open Web
Web agents--autonomous systems that navigate and execute tasks on the web on behalf of users--have the potential to transform how people interact with the digital world. However, the most capable…
XYBench: Can LLMs Respond Pragmatically to Queries with Misconceptions?
When non-expert users ask LLMs for assistance, their queries can often have misconceptions (e.g.,"How do I parse XML with regex?"). In such cases, often referred to as the XY-problem, LLMs must…
Measuring AI Scientists: From Exams to Discovery
Large language models and agentic systems are increasingly embedded across the scientific work-flow, from literature synthesis and hypothesis generation to code execution, data analysis and writing.…
Process-Oriented Evaluation of AI-Assisted Scientific Writing
Bad writing hinders the publication of science. The role of artificial intelligence (AI) in generating and editing scientific texts remains unsettled. Abstracts serve as the critical gateway to…
Stochastic Emulation of a Fully Coupled Preindustrial E3SMv3 Simulation
We present a stochastic coupled emulator of E3SM version 3, built on the SamudrACE framework, which couples an atmosphere emulator (ACE2) with a full-depth ocean emulator (Samudra). We replace the…
ArtifactLinker: Linking Scientific Artifacts for Automatic State-of-the-Art Discovery
Scientific artifacts such as models and datasets are foundations for research. With the rapid growth of platforms like HuggingFace, researchers now have access to a large number of artifacts. Yet, a…
CoTs as Tractable Probabilistic Programs
Chain-of-thought (CoT) traces are used across language model prompting, training, test-time inference, and interpretability, yet they are often modeled in task-specific ways. We propose treating CoT…
Querying Multimodal Scientific Papers with AI: Practices and Preferences Across Blind, Low-Vision, and Sighted Scientists
Visual diagrams, figures, and tables are central to scientific papers, and convey information beyond what is captured in text. While blind or low-vision (BLV) scientists have traditionally relied on…
Rethinking the Evaluation of Harness Evolution for Agents
We revisit the evaluation of automatic harness evolution for LLM agents. Existing harness evolution methods use unit test cases to search for harness configurations and then report final performance…