Skip to main content ->
Ai2

How researchers adapted Dolma for better Thai language models

August 26, 2026

Ai2


When we release open tools, we do so with the idea that researchers will take them in directions we couldn’t anticipate. We want them to build on our work, test new ideas, and apply those approaches to new problems and communities.

That’s why we created Dolma as an open toolkit for curating the large-scale text datasets used to train language models, including our open Olmo models. Instead of releasing only finished datasets, we made the data-curation process available for researchers to study, adapt, and tailor for their own needs.

The Mangosteen project shows what that can look like in practice. A team of Thai researchers used Dolma as the foundation for a 47-billion-token Thai pretraining corpus, adapting the pipeline for Thai data, where existing approaches didn’t always account for the language’s unique characteristics.

The team used that flexibility to address a challenge in building high-quality training datasets for Thai language models. Publicly available Thai pretraining datasets had not been extensively audited by Thai speakers and were often built largely from web crawls—large collections of text automatically gathered from websites across the internet. They found that some existing datasets contained content they considered unsuitable for training, while missing valuable sources of Thai text, including books, research papers, official websites, and YouTube subtitles.

For a small research team, building a complete data-curation pipeline would’ve been a major undertaking. “We do not have enough developers to build a data-curation pipeline from scratch,” says Wannaphong Phatthiyaphaibun, one of the project leads and a PhD student at Vidyasirimedhi Institute of Science and Technology (VISTEC). Dolma gave the researchers a strong starting point for turning raw text into training data, allowing them to focus on improving the process for Thai rather than rebuilding the entire workflow themselves.

When the researchers adapted Dolma for Thai, they found that some parts of the pipeline needed to be redesigned. Sentence- and paragraph-level deduplication, for example, removed almost all of their data because Thai doesn’t define sentence boundaries in the same way as English. The team kept the pipeline components that worked, such as document- and URL-level deduplication, while modifying other steps for Thai. 

They also tweaked certain quality filters, replaced language-specific tools, and added rules for patterns common in Thai web data. For example, they found that some Thai news pages contained only truncated snippets followed by “Read More” prompts, requiring additional filtering to remove incomplete articles. 

The team’s adaptations produced Mangosteen, a corpus designed around the needs of Thai language modeling. In experiments with Thai LLMs, they found that their curation pipeline could remove more than 80% of the Common Crawl data they started with and nearly half of FineWeb2 – two large collections of web data, the latter already cleaned and curated for model training – while maintaining or improving performance compared with models trained on the larger datasets. 

Those improvements carried into bigger models as well. Models trained on Mangosteen showed stronger performance on Thai cultural knowledge evaluations, which the team sees as evidence that locally relevant data can help models better represent the communities they serve. 

“The open-source nature of Dolma allows us to inspect, modify, test, and reproduce the results for our language,” says Phatthiyaphaibun. He added that openness allowed the team to use their own expertise to create training data “that serves our local community while contributing to global LLM development.” 

Mangosteen is one example of what can happen when researchers have access not only to models and datasets, but also to the tools that shape them. By making those foundations open, we hope to help researchers around the world adapt AI systems for the languages and communities they understand best.

Join us

At Ai2 we’re building the future of transparent, open-source AI — built in the open to empower scientific progress and fundamental understanding of this world changing technology. We’re not here to make profits, we’re here to make sure benefits of AI are shared widely and for the benefit of humanity. If this appeals to you, please take a look at our open roles.

Subscribe to receive monthly updates about the latest Ai2 news.