An abstract illustration of swirling shapes, meant to denote a futuristic feeling.

Research - Papers

Explore a selection of our published work on a variety of key research challenges in AI.

Elaboration-Generating Commonsense Question Answering at Scale

Wenya WangVivek SrikumarHannaneh HajishirziNoah A. Smith

2023

ACL

In question answering requiring common sense, language models (e.g., GPT-3) have been used to generate text expressing background knowledge that helps improve performance. Yet the cost of working…

Few-shot Fine-tuning vs. In-context Learning: A Fair Comparison and Evaluation

Marius MosbachTiago PimentelShauli RavfogelYanai Elazar

2023

Findings of ACL 2023

Few-shot fine-tuning and in-context learning are two alternative strategies for task adaptation of pre-trained language models. Recently, in-context learning has gained popularity over fine-tuning…

FiD-ICL: A Fusion-in-Decoder Approach for Efficient In-Context Learning

Qinyuan YeIz BeltagyMatthew E. PetersHannaneh Hajishirzi

2023

ACL

Large pre-trained models are capable of few-shot in-context learning (ICL), i.e., performing a new task by prepending a few demonstrations before the test input. However, the concatenated…

HINT: Hypernetwork Instruction Tuning for Efficient Zero-Shot Generalisation

Hamish IvisonAkshita BhagiaYizhong WangMatthew E. Peters

2023

ACL

Recent NLP models have the great ability to generalise ‘zero-shot’ to new tasks using only an instruction as guidance. However, these approaches usually repeat their instructions with every input,…

NarrowBERT: Accelerating Masked Language Model Pretraining and Inference

Haoxin LiPhillip KeungDaniel ChengNoah A. Smith

2023

ACL • Proceedings

Large-scale language model pretraining is a very successful form of self-supervised learning in natural language processing, but it is increasingly expensive to perform as the models and pretraining…

Nonparametric Masked Language Modeling

Sewon MinWeijia ShiM. LewisLuke Zettlemoyer

2023

ACL • Findings

Existing language models (LMs) predict tokens with a softmax over a finite vocabulary, which can make it difficult to predict rare tokens or phrases. We introduce NPM, the first nonparametric masked…

One Embedder, Any Task: Instruction-Finetuned Text Embeddings

Hongjin SuWeijia ShiJungo KasaiTao Yu

2023

ACL • Findings

We introduce INSTRUCTOR, a new method for computing text embeddings given task instructions: every text input is embedded together with instructions explaining the use case (e.g., task and domain…

PuMer: Pruning and Merging Tokens for Efficient Vision Language Models

Qingqing CaoBhargavi ParanjapeHanna Hajishirzi

2023

ACL

Large-scale vision language (VL) models use Transformers to perform cross-modal interactions between the input text and image. These cross-modal interactions are computationally expensive and…

Risks and NLP Design: A Case Study on Procedural Document QA

Nikita HaduongAlice GaoNoah A. Smith

2023

ACL • Findings

As NLP systems are increasingly deployed at scale, concerns about their potential negative impacts have attracted the attention of the research community, yet discussions of risk have mostly been at…

Self-Instruct: Aligning Language Models with Self-Generated Instructions

Yizhong WangYeganeh KordiSwaroop MishraHannaneh Hajishirzi

2023

ACL

Large “instruction-tuned” language models (i.e., finetuned to respond to instructions) have demonstrated a remarkable ability to generalize zero-shot to new tasks. Nevertheless, they depend heavily…

Previous82-91Next