Skip to main content ->
Ai2

Research - Papers

Explore a selection of our published work on a variety of key research challenges in AI.

Filter papers

Process-Oriented Evaluation of AI-Assisted Scientific Writing

Patrick Queiroz Da SilvaSanchaita HazraDoeun LeeBodhisattwa Prasad Majumder
2026
Conference on Language Modeling

Bad writing hinders the publication of science. The role of artificial intelligence (AI) in generating and editing scientific texts remains unsettled. Abstracts serve as the critical gateway to… 

ArtifactLinker: Linking Scientific Artifacts for Automatic State-of-the-Art Discovery

Haofei YuJiaxuan YouPeter ClarkKyle Richardson
2026
COLM 2026

Scientific artifacts such as models and datasets are foundations for research. With the rapid growth of platforms like HuggingFace, researchers now have access to a large number of artifacts. Yet, a… 

CoTs as Tractable Probabilistic Programs

Kyle RichardsonYu FengPoorva GargDan Roth
2026
9th Workshop on Tractable Probabilistic Modeling (TPM@UAI)

Chain-of-thought (CoT) traces are used across language model prompting, training, test-time inference, and interpretability, yet they are often modeled in task-specific ways. We propose treating CoT… 

Evidence-Informed LLM Beliefs for Continual Scientific Discovery

Dhruv AgarwalReece AdamsonAndrew McCallumBodhisattwa Prasad Majumder
2026
arXiv.org

Open-ended scientific discovery with large language models (LLMs) increasingly operates as a long-horizon loop of hypothesis search and verification, where a reward signal guides which hypotheses to… 

Operadic consistency: a label-free signal for compositional reasoning failures in LLMs

Nathaniel BottmanYinhong LiuKyle Richardson
2026
arXiv

Detecting LLM reasoning failures at inference time without ground-truth labels has motivated a wide range of confidence baselines, including self-consistency, semantic entropy, and P(True), built on… 

Operads for compositional reasoning in LLMs

Nathaniel BottmanKyle Richardson
2026
Workshop on Combining Theory and Benchmarks (@ICML 2026)

Question decomposition, i.e. breaking a complex query into simpler sub-queries whose answers are composed to produce a final answer, is a widely used strategy for improving LLM reasoning, yet it… 

Scientific reasoning does not reliably translate into scientific forecasting in frontier AI

Sean WuPan LuYupeng ChenJunchi Yu
2026
arXiv

AI systems are increasingly used to support forward-looking scientific judgment, but it remains unclear whether they can form reliable expectations about future scientific advances. Here we show… 

Analytica: Soft Propositional Reasoning for Robust and Scalable LLM-Driven Analysis

Junyan ChengKyle RichardsonPeter Chin
2026
ICLR 2026

Large language model (LLM) agents are increasingly tasked with complex real-world analysis (e.g., in financial forecasting, scientific discovery), yet their reasoning suffers from stochastic… 

Probabilistic Programs of Thought

Poorva GargRenato Lui GehDaniel Mingyi IsraelGuy Van den Broeck
2026
arXiv

LMs are widely used for code generation and mathematical reasoning tasks where they are required to generate structured output. They either need to reason about code, generate code for a given… 

Leveraging In-Context Learning for Language Model Agents

Shivanshu GuptaSameer SinghAshish SabharwalBen Bogin
2025
NeurIPS • Workshop on Multi-Turn Interactions in LLMs

In-context learning (ICL) with dynamically selected demonstrations combines the flexibility of prompting large language models (LLMs) with the ability to leverage training data to improve… 

1-10Next