Research - Papers

Explore a selection of our published work on a variety of key research challenges in AI.

Measuring and Improving Attentiveness to Partial Inputs with Counterfactuals

Yanai ElazarBhargavi ParanjapeHao PengNoah A. Smith

2024

EMNLP Findings

The inevitable appearance of spurious correlations in training datasets hurts the generalization of NLP models on unseen data. Previous work has found that datasets with paired inputs are prone to…

Mechanistic?

Naomi SaphraSarah Wiegreffe

2024

EMNLP • BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP

The rise of the term “mechanistic interpretability” has accompanied increasing interest in understanding neural models—particularly language models. However, this jargon has also led to a fair…

Merge to Learn: Efficiently Adding Skills to Language Models with Model Merging

Jacob Daniel MorrisonNoah A. SmithHanna HajishirziPradeep Dasigi

2024

EMNLP Findings

Adapting general-purpose language models to new skills is currently an expensive process that must be repeated as new instruction datasets targeting new skills are created, or can cause the models…

Plausibly Problematic Questions in Multiple-Choice Benchmarks for Commonsense Reasoning

Shramay PaltaNishant BalepurPeter RankelRachel Rudinger

2024

EMNLP Findings

Questions involving commonsense reasoning about everyday situations often admit many possible or plausible answers. In contrast, multiple-choice question (MCQ) benchmarks for commonsense reasoning…

Scalable Data Ablation Approximations for Language Models through Modular Training and Merging

Clara NaIan MagnussonAnanya Harsh JhaPradeep Dasigi

2024

EMNLP

Training data compositions for Large Language Models (LLMs) can significantly affect their downstream performance. However, a thorough data ablation study exploring large sets of candidate data…

SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research Repositories

Ben BoginKejuan YangShashank GuptaTushar Khot

2024

EMNLP

Given that Large Language Models (LLMs) have made significant progress in writing code, can they now be used to autonomously reproduce results from research repositories? Such a capability would be…

ComPO: Community Preferences for Language Model Personalization

Sachin KumarChan Young ParkYulia TsvetkovHanna Hajishirzi

2024

arXiv.org

Conventional algorithms for training language models (LMs) with human feedback rely on preferences that are assumed to account for an"average"user, disregarding subjectivity and finer-grained…

CLIN: A Continually Learning Language Agent for Rapid Task Adaptation and Generalization

Bodhisattwa Prasad MajumderBhavana Dalvi MishraPeter JansenPeter Clark

2024

COLM

Language agents have shown some ability to interact with an external environment, e.g., a virtual world such as ScienceWorld, to perform complex tasks, e.g., growing a plant, without the startup…

SynerGPT: In-Context Learning for Personalized Drug Synergy Prediction and Drug Design

Carl N. EdwardsAakanksha NaikTushar KhotTom Hope

2024

COLM

Predicting synergistic drug combinations can help accelerate discovery of cancer treatments, particularly therapies personalized to a patient’s specific tumor via biopsied cells. In this paper, we…

m&m's: A Benchmark to Evaluate Tool-Use for multi-step multi-modal Tasks

Zixian MaWeikai HuangJieyu ZhangRanjay Krishna

2024

ECCV

Real-world multi-modal problems are rarely solved by a single machine learning model, and often require multi-step computational plans that involve stitching several models. Tool-augmented LLMs hold…

Previous81-90Next