Skip to main content ->
Ai2

Latest research

September 1, 2026

BenchMIRT: What are LLM benchmarks actually measuring?

BenchMIRT is a new method for auditing LLM benchmarks question by question, revealing which capabilities they actually measure and helping researchers build smaller, more focused, and easier-to-interpret evaluations.
Read post
August 7, 2026

TutorMoments: Do AI tutors know when to help and when to hold back?

TutorMoments is an open, replay-based evaluation framework that tests whether AI tutors can recognize when to support a student and when to hold back and encourage deeper reasoning.
Read post
July 28, 2026

The OlmoEarth Platform: Geospatial inference at planetary scale

How we built the OlmoEarth Platform to fine-tune geospatial models and run continent-scale satellite inference while managing massive data pipelines, distributed compute, and automatically recovering from failures at scale.
Read post
July 13, 2026

What building Shippy taught us about building agents

Building Shippy taught us that reliable agents depend less on the model itself than on deterministic tools, explicit guardrails, isolated infrastructure, and evaluations grounded in real-world workflows and live data.
Read post
June 29, 2026

DiScoFormer: One transformer for density and score, across distributions

DiScoFormer is a transformer-based density and score estimator that can infer both quantities from a finite sample in one forward pass, generalizing classical KDE while staying accurate in high-dimensional and out-of-distribution settings without retraining for each new distribution.
Read post
June 25, 2026

Which tokens does a hybrid model predict better?

New token-level analyses of Olmo 3 and Olmo Hybrid show that hybrid models predict meaning-bearing, context-dependent tokens better than transformers, while transformers retain an edge on verbatim copying.
Read post
June 17, 2026

MolmoMotion: Language-guided 3D motion forecasting

MolmoMotion is an open, language-guided 3D motion forecasting model that predicts how object points will move in the future, enabling stronger motion prediction for robotics, video generation, and other systems that need to reason about what happens next.
Read post
June 12, 2026

olmo-eval: An evaluation workbench for the model development loop

olmo-eval is an open evaluation workbench that helps model developers add, run, and analyze benchmarks across changing LLM checkpoints, extending OLMES from final-score reproducibility into the day-to-day model development loop.
Read post
May 19, 2026

OlmoEarth v1.1: A more efficient family of models

OlmoEarth v1.1 is a more efficient family of remote-sensing models that cuts compute costs by up to 3x while maintaining similar performance, making large-scale satellite mapping faster and cheaper to run.
Read post
1-9Next