Research - Papers
Explore a selection of our published work on a variety of key research challenges in AI.
AIMIP Phase 1: systematic evaluations of AI weather and climate models
We present the AI weather and climate model intercomparison project (AIMIP), phase 1. Drawing from the rich tradition of intercomparisons in climate model development, we specify a common…
Improving Attributed Long-form Question Answering with Intent Awareness
Large language models (LLMs) are increasingly being used to generate comprehensive, knowledge-intensive reports. However, while these models are trained on diverse academic papers and reports, they…
DRACULA: Hunting for the Actions Users Want Deep Research Agents to Execute
Scientific Deep Research (DR) agents answer user queries by synthesizing research papers into multi-section reports. User feedback can improve their utility, but existing protocols only score the…
Analytica: Soft Propositional Reasoning for Robust and Scalable LLM-Driven Analysis
Large language model (LLM) agents are increasingly tasked with complex real-world analysis (e.g., in financial forecasting, scientific discovery), yet their reasoning suffers from stochastic…
AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite
AI agents hold great real-world promise, with the potential to revolutionize scientific productivity by automating literature reviews, replicating experiments, analyzing data, and even proposing new…
Improving Attributed Long-form Question Answering with Intent Awareness
Large language models (LLMs) are increasingly being used to generate comprehensive, knowledge-intensive reports. However, while these models are trained on diverse academic papers and reports, they…
On the Reasoning Abilities of Masked Diffusion Language Models
Masked diffusion models (MDMs) for text offer a compelling alternative to traditional autoregressive language models. Parallel generation makes them efficient, but their computational capabilities…
SimpleToM: Exposing the Gap between Explicit ToM Inference and Implicit ToM Application in LLMs
Large language models (LLMs) are increasingly tested for a"Theory of Mind"(ToM) - the ability to attribute mental states to oneself and others. Yet most evaluations stop at explicit belief…
Probabilistic Programs of Thought
LLMs are widely used for code generation and mathematical reasoning tasks where they are required to generate structured output. They either need to reason about code, generate code for a given…
Cocoa: Co-Planning and Co-Execution with AI Agents
As AI agents take on increasingly long-running tasks involving sophisticated planning and execution, there is a corresponding need for novel interaction designs that enable deeper human-agent…