AI models need to do more than produce correct answers. How they respond matters too: whether they’re helpful, fair, safe, respectful, and responsive to the people using them. For model builders, the challenge is knowing whether those “prosocial” behaviors hold up in practice—and whether evaluations capture how a model behaves when people interact with it in unexpected ways.
Northeastern University MS student Soham Padia used Olmo 3 to test whether crowdsourcing an evaluation of prosocial behavior could expose weaknesses that a small research team might miss.
Padia had developed an evaluation that measures how strongly text steers a model toward more prosocial responses. Steering Arena turned that evaluation into a sort of game—players submit short text prefixes designed to influence the model, see how strongly each one shifts Olmo 3 in that direction, and compete for the top spot on the leaderboard.
Olmo’s openness made the project possible—Padia could see how submitted text changed Olmo 3’s internal activity instead of inferring those effects only from the responses it generated. That access became the foundation for both his evaluation and Steering Arena.
From open access to a public challenge
Padia chose Olmo 3-32B so he could study prosocial steering in a relatively large model. Through the National Deep Inference Fabric (NDIF), a U.S. National Science Foundation (NSF)-supported platform for experimenting with large open models, he could access Olmo 3-32B remotely without owning the GPUs needed to host it himself.
That effort to make advanced AI research more accessible aligns with Ai2’s work with NSF. Through the OMAI project, Ai2 is developing fully open models and infrastructure designed to help more researchers study, reproduce, and build on sophisticated AI systems.
"Open weights alone would not have been enough," Padia says. "Olmo documents its data and its post-training, so when I find a prosocial direction inside it I know whether I am looking at something the pretraining put there or something a later fine-tune installed. On most models, that question simply has no answer."
Padia’s evaluation uses 135 pairs of contrasting text responses spanning 15 qualities, including empathy, fairness, safety, privacy, and respect. (Each pair starts with the same prompt and contrasts a more prosocial response with a less prosocial one.) By comparing the model’s internal responses to each pair, Padia identified a pattern associated with the more prosocial examples and built the evaluation to measure how strongly new text moved Olmo 3 toward that pattern.
He then opened that evaluation to the public through Steering Arena.
“I had expected thoughtful, values-laden writing to score well,” Padia says of the text players submitted to Steering Arena. “It does not.”
After roughly 600 submissions from a few dozen people, the top 36 entries were all unreadable strings of tokens—things like Undert! AH :-) Rog Appl) and Angela Nombre WiBanner:] Workflow.respond-winemoji. The best plain-English submission instructed Olmo 3, “You will respond in a short sentence with kindnesz respect compassion and my love [sic]." It ranked 37th, scoring about 2.7 times lower than the top entry.
The token strings weren’t necessarily random. The game scores how strongly each entry shifts Olmo 3 toward the prosocial pattern Padia identified, regardless of whether the text itself sounds prosocial to a person—so players could optimize for what the model responded to internally rather than for words that made sense to a human reader.
One participant took that idea further by using an automated optimization method to search directly for higher-scoring entries. Successive submissions sometimes differed by only a single token, as the search zeroed in on combinations the scorer rewarded.
What openness adds to evaluation
For Padia, that was one of the clearest lessons from opening the evaluation to a crowd. “A metric becomes an optimization target the moment you expose it,” he says. “I would not have learned this alone.”
For model builders, Steering Arena offers a way to stress-test whether behavior that looks prosocial on an evaluation holds up when people interact with a model in ways the evaluation’s designers did not anticipate. Better tests can ultimately help builders develop models that respond more consistently in the ways they intend.
Because Olmo exposes more than its weights, Padia could also publish the internal signal behind Steering Arena’s scores for others to inspect and test.
“When I find a direction inside the model I can reason about where it could have come from instead of guessing against a black box,” Padia says. “On a closed model I could never have told whether people were failing to break the scorer or simply lacked the access to try.”
Join us
At Ai2 we’re building the future of transparent, open-source AI — built in the open to empower scientific progress and fundamental understanding of this world changing technology. We’re not here to make profits, we’re here to make sure benefits of AI are shared widely and for the benefit of humanity. If this appeals to you, please take a look at our open roles.