Latest research
August 19, 2025
Signal and Noise: Reducing uncertainty in language model evaluation
We find that two simple metrics, signal and noise, reveal key differences in the utility of current LLM benchmarks.August 18, 2025
MoNaCo: More natural questions for reasoning across dozens of documents
Introducing MoNaCo, a benchmark of highly challenging questions spanning dozens of documents for evaluating large language models.August 12, 2025
MolmoAct: An Action Reasoning Model that reasons in 3D space
MolmoAct is the first model able to “think” in three dimensions, trained efficiently and delivering benchmark-topping performance.July 22, 2025
Contextualized Evaluations: Judging language model responses to underspecified queries
How do we evaluate LLMs on underspecified queries? We show that adding clarifying context flips model rankings and uncovers model biases.July 18, 2025
AutoDS: A prototype engine for autonomous, open-ended scientific discovery
AutoDS goes beyond standard data crunching by building upon its own findings and uncovering insights that may not be immediately apparent even to experienced researchers.July 9, 2025
Introducing FlexOlmo: a new paradigm for language model training and data collaboration
Explore how FlexOlmo enables collaborative language model training without sacrificing data privacy or control, introducing a new, flexible approach to building shared AI models.July 1, 2025
SciArena: A new platform for evaluating foundation models in scientific literature tasks
Discover how SciArena is being used to evaluate foundation models’ capabilities in scientific literature tasks through community-driven, literature-grounded, and multi-disciplinary reasoning.June 24, 2025
OMEGA: Can LLMs reason outside the box in math?
Discover how OMEGA is being used to evaluate large language models' ability to generalize in math through exploratory, compositional, and transformative reasoningJune 13, 2025