Skip to main content ->
Ai2

Research - Papers

Explore a selection of our published work on a variety of key research challenges in AI.

Filter papers

OLMES: A Standard for Language Model Evaluations

Yuling GuOyvind TafjordBailey KuehlHanna Hajishirzi
2024
arXiv.org

Progress in AI is often demonstrated by new models claiming improved performance on tasks measuring model capabilities. Evaluating language models in particular is challenging, as small changes to… 

SelfGoal: Your Language Agents Already Know How to Achieve High-level Goals

Ruihan YangJiangjie ChenYikai ZhangDeqing Yang
2024
technical report

Language agents powered by large language models (LLMs) are increasingly valuable as decision-making tools in domains such as gaming and programming. However, these agents often face challenges in… 

Digital Socrates: Evaluating LLMs through explanation critiques

Yuling GuOyvind TafjordPeter Clark
2024
ACL

While LLMs can provide reasoned explanations along with their answers, the nature and quality of those explanations are still poorly understood. In response, our goal is to define a detailed way of… 

Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMs

Shashank GuptaVaishnavi ShrivastavaA. DeshpandeTushar Khot
2024
ICLR

Recent works have showcased the ability of LLMs to embody diverse personas in their responses, exemplified by prompts like 'You are Yoda. Explain the Theory of Relativity.' While this ability allows… 

The Expressive Power of Transformers with Chain of Thought

William MerrillAshish Sabharwal
2024
ICLR

Recent theoretical work has identified surprisingly simple reasoning problems, such as checking if two nodes in a graph are connected or simulating finite-state machines, that are provably… 

Closing the Curious Case of Neural Text Degeneration

Matthew FinlaysonJohn HewittAlexander KollerAshish Sabharwal
2024
ICLR

Despite their ubiquity in language generation, it remains unknown why truncation sampling heuristics like nucleus sampling are so effective. We provide a theoretical explanation for the… 

Enhancing Systematic Decompositional Natural Language Inference Using Informal Logic

Nathaniel WeirKate SandersOrion WellerBenjamin Van Durme
2024
arXiv.org

Contemporary language models enable new opportunities for structured reasoning with text, such as the construction and evaluation of intuitive, proof-like textual entailment trees without relying on… 

Calibrating Large Language Models with Sample Consistency

Qing LyuKumar ShridharChaitanya MalaviyaChris Callison-Burch
2024
arXiv

Accurately gauging the confidence level of Large Language Models' (LLMs) predictions is pivotal for their reliable application. However, LLMs are often uncalibrated inherently and elude conventional… 

TimeArena: Shaping Efficient Multitasking Language Agents in a Time-Aware Simulation

Yikai ZhangSiyu YuanCaiyu HuJiangjie Chen
2024
ACL 2024

Despite remarkable advancements in emulating human-like behavior through Large Language Models (LLMs), current textual simulations do not adequately address the notion of time. To this end, we… 

OLMo: Accelerating the Science of Language Models

Dirk GroeneveldIz BeltagyPete WalshHanna Hajishirzi
2024
ACL 2024

Language models (LMs) have become ubiquitous in both NLP research and in commercial product offerings. As their commercial importance has surged, the most powerful models have become closed off,…