Skip to main content ->
Ai2

Latest research

October 7, 2026

Now in Nature: Retrofitting language models to operate over bytes

The technique behind Bolmo, Ai2’s fully open byte-level language models, is now published in Nature, with new checkpoints showing the approach generalizes beyond Olmo to other model families.
Read post
October 2, 2026

Open-sourcing AstaBrief, the fast report-generation model in Asta

We’re releasing AstaBrief, an 8B open-weights model for generating cited scientific reports, available in Asta’s Fast mode or to download and run on your own infrastructure.
Read post
October 1, 2026

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

Olmo-core 3 introduces a redesigned, fully open training stack for efficiently scaling mixture-of-experts models into the trillion-parameter range.
Read post
September 1, 2026

BenchMIRT: What are LLM benchmarks actually measuring?

BenchMIRT is a new method for auditing LLM benchmarks question by question, revealing which capabilities they actually measure and helping researchers build smaller, more focused, and easier-to-interpret evaluations.
Read post
August 7, 2026

TutorMoments: Do AI tutors know when to help and when to hold back?

TutorMoments is an open, replay-based evaluation framework that tests whether AI tutors can recognize when to support a student and when to hold back and encourage deeper reasoning.
Read post
July 28, 2026

The OlmoEarth Platform: Geospatial inference at planetary scale

How we built the OlmoEarth Platform to fine-tune geospatial models and run continent-scale satellite inference while managing massive data pipelines, distributed compute, and automatically recovering from failures at scale.
Read post
July 13, 2026

What building Shippy taught us about building agents

Building Shippy taught us that reliable agents depend less on the model itself than on deterministic tools, explicit guardrails, isolated infrastructure, and evaluations grounded in real-world workflows and live data.
Read post
June 29, 2026

DiScoFormer: One transformer for density and score, across distributions

DiScoFormer is a transformer-based density and score estimator that can infer both quantities from a finite sample in one forward pass, generalizing classical KDE while staying accurate in high-dimensional and out-of-distribution settings without retraining for each new distribution.
Read post
June 25, 2026

Which tokens does a hybrid model predict better?

New token-level analyses of Olmo 3 and Olmo Hybrid show that hybrid models predict meaning-bearing, context-dependent tokens better than transformers, while transformers retain an edge on verbatim copying.
Read post
1-9Next