Latest research
October 7, 2026
Now in Nature: Retrofitting language models to operate over bytes
The technique behind Bolmo, Ai2’s fully open byte-level language models, is now published in Nature, with new checkpoints showing the approach generalizes beyond Olmo to other model families.October 2, 2026
Open-sourcing AstaBrief, the fast report-generation model in Asta
We’re releasing AstaBrief, an 8B open-weights model for generating cited scientific reports, available in Asta’s Fast mode or to download and run on your own infrastructure.October 1, 2026
Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs
Olmo-core 3 introduces a redesigned, fully open training stack for efficiently scaling mixture-of-experts models into the trillion-parameter range.September 1, 2026
BenchMIRT: What are LLM benchmarks actually measuring?
BenchMIRT is a new method for auditing LLM benchmarks question by question, revealing which capabilities they actually measure and helping researchers build smaller, more focused, and easier-to-interpret evaluations.August 7, 2026
TutorMoments: Do AI tutors know when to help and when to hold back?
TutorMoments is an open, replay-based evaluation framework that tests whether AI tutors can recognize when to support a student and when to hold back and encourage deeper reasoning.July 28, 2026
The OlmoEarth Platform: Geospatial inference at planetary scale
How we built the OlmoEarth Platform to fine-tune geospatial models and run continent-scale satellite inference while managing massive data pipelines, distributed compute, and automatically recovering from failures at scale.July 13, 2026
What building Shippy taught us about building agents
Building Shippy taught us that reliable agents depend less on the model itself than on deterministic tools, explicit guardrails, isolated infrastructure, and evaluations grounded in real-world workflows and live data.June 29, 2026
DiScoFormer: One transformer for density and score, across distributions
DiScoFormer is a transformer-based density and score estimator that can infer both quantities from a finite sample in one forward pass, generalizing classical KDE while staying accurate in high-dimensional and out-of-distribution settings without retraining for each new distribution.June 25, 2026
Which tokens does a hybrid model predict better?
New token-level analyses of Olmo 3 and Olmo Hybrid show that hybrid models predict meaning-bearing, context-dependent tokens better than transformers, while transformers retain an edge on verbatim copying.1-9Next