Skip to main content ->
Ai2

Ai2 Newsletter

August 2026

Top story - Ai2 and Hugging Face expand their collaboration for truly open AI

Today we're expanding our partnership with Hugging Face to make our fully open models and datasets easier to find, customize, and build on.

The Hugging Face Hub is a central platform for the AI community, and already where many people go to find and use our open work. Under the new agreement, Hugging Face is roughly tripling our storage there to nearly two petabytes, effective immediately, and our traffic is no longer subject to the Hub's standard rate limits—so even our largest datasets and multi-checkpoint models download at full throughput.

Releasing model weights is only one part of our continued open science effort. We also publish the training data, intermediate checkpoints, data mixes, evaluations, ablations, and demos that let other researchers reproduce our work and examine how our artifacts were developed.

According to Hugging Face's official heatmap, we publish more new artifacts each year than any other organization it tracks, and the added capacity is what keeps all of those artifacts easily obtainable as our portfolio grows—including forthcoming work on multimodal models, tool-using agents, and open models for scientific communities.

Hugging Face is also a collaborator on our open science itself.

Its team wired olmOCR-Bench, our benchmark for how well models read real documents, into the Hub's leaderboard workflow. When it brought MolmoAct 2 into LeRobot, its open-source robotics library, we released the model's training data in LeRobot's standard format so developers could run it on Hugging Face's low-cost SO-100 and SO-101 robotic arms. Benchmarks like IFBench, RewardBench 2, and AstaBench live on the Hub too, alongside our OlmoEarth models for Earth observation and climate models like HiRO-ACE—both a part of Hugging Face's Hugging Science collection.

Together with Hugging Face, we're working toward an ecosystem where anyone can build with, inspect, and run experiments using a wide range of AI tools. This agreement is a major step in that direction.

Two research-focused updates to Asta

Find papers, Asta's literature search tool, now runs Deep search by default. Deep search is an upgraded engine that interprets your query, evaluates whether the results answer it, and re-searches where they fall short. And when an AutoDiscovery hypothesis is worth a closer look, a single "Explore with Asta" click opens it in DataVoyager with your datasets and results already loaded—no need to set it up again.

Geospatial inference at planetary scale

A new engineering write-up covers the infrastructure behind our open Earth observation models, from splitting an inference job across high-I/O CPUs and GPUs to aligning imagery from providers that use different projections. The platform can run inference across a continent in roughly a day, at fractions of a penny per square kilometer.

Tracing distinctive language in AI-written text

Tuhin Chakrabarty, an assistant professor at Stony Brook University, and his group are using our infini-gram engine to trace where the language in books flagged as AI-written comes from. Because infini-gram indexes massive public text datasets and counts how often a phrase of any length appears across them, it can show which expressions a passage shares with earlier writing, and where.

What building our Shippy agent taught us

We published a deep dive on Shippy, the agentic assistant our Skylight team built for real-time maritime domain awareness. An analyst can ask it something like, "Show me fishing activity in Panama's EEZ last month," and Shippy resolves the country's exclusive economic zone to a boundary, queries live vessel and satellite data, and links back to the Skylight map so every number and detail can be checked.

Building on FlexOlmo

Researchers at the University of Southern Denmark took our FlexOlmo architecture and extended it for Danish Foundation Models, a national project building open models for the Danish language. Their version, FlexMoRE, shrinks most of FlexOlmo's full-size components into much smaller low-rank adapters, so a modular model can run on consumer hardware.

olmOCR 2 in the Ai2 Playground

olmOCR 2, our model for turning real documents into structured text, is now in the Ai2 Playground, so you can try it on your own files without setting up a bespoke pipeline. It does well on the cases that usually break OCR: handwriting, dense tables, and complex multi-column layouts.

OlmoEarth at the Global Nature Positive Summit

Ted Schmitt, our Senior Director of Conservation, brought OlmoEarth to the second Global Nature Positive Summit in Kumamoto, Japan – which opened July 14 and drew 2,725 delegates – together with our partners at the Group on Earth Observations. His session covered the data needed to measure the state of ecosystems, less than 20% of which are under satellite monitoring today.

Ai2 at Seattle Tech Week

On July 30, during Seattle Tech Week, we opened our Northlake office for an afternoon on how AI actually gets built at Ai2, and 147 people turned up. Iz Beltagy, our Olmo research lead, walked through the technical work behind our open models—what it takes to scale a large training run, and how fine-tuning, reinforcement learning, post-training, and evaluation fit together before anything gets released.

Thank you to everyone who came out, and to those who stayed after to talk about what they're working on.

    Ai2 Newsletter Archive