Event
LMSS @ Cornell Tech: Swabha Swayamdipta (USC)
Learning Machines Seminar Series What: LMSS: Swabha Swayamdipta (USC) When: Thursday, October 16, 1:30-2:45 pm Where: Bloomberg 081, Bloomberg Center, Cornell Tech (map) The series is organized by Associate Professor Yoav Artzi and sponsored by Bloomberg. Pizza will be served at 1:15 p.m. "Safety and Accountability for Language Models with Simple Learning Recipes" While widely recognized as important goals, safety and accountability in language models remain elusive. In this talk, I will highlight some recent work where well-known fundamental learning recipes, such as the cross entropy objective, can be shown to improve language model safety and accountability. First, I will discuss a selective loss function that teaches language models, during pretraining, to understand but not generate high-risk data. Second, I will present a sequence-to-sequence learning method for prompt inversion from logprob sequences that recovers hidden prompts by gleaning clues from the model’s next-token probabilities over the course of multiple generation steps. Together, these methods highlight how simple learning recipes can complement elaborate and resource intensive post-training approaches for safety and accountability. BIO Swabha Swayamdipta is an Assistant Professor of Computer Science and a co-Associate Director of the Center for AI and Society at the University of Southern California. Her research interests are in natural language processing and machine learning, with a primary interest in the evaluation of generative models of language, understanding the behavior of language models, and designing language technologies for societal good. At USC, Swabha leads the Data, Interpretability, Language and Learning (DILL) Lab. She received her PhD from Carnegie Mellon University, followed by a postdoc at the Allen Institute for AI and the University of Washington. Her work has received outstanding paper awards at EMNLP 2024, ICML 2022, NeurIPS 2021 and ACL 2020. Her research is supported by awards from NSF, Apple, the Allen Institute for AI, Intel Labs, the Zumberge Foundation and a WiSE Gabilan Fellowship.