Skip to content

[Paper Note] Foundational Challenges in Assuring Alignment and Safety of Large Language Models, Usman Anwar+, TMLR'24 #1868

Description

@AkihikoWatanabe

URL

Authors

  • Usman Anwar
  • Abulhair Saparov
  • Javier Rando
  • Daniel Paleka
  • Miles Turpin
  • Peter Hase
  • Ekdeep Singh Lubana
  • Erik Jenner
  • Stephen Casper
  • Oliver Sourbut
  • Benjamin L. Edelman
  • Zhaowei Zhang
  • Mario Günther
  • Anton Korinek
  • Jose Hernandez-Orallo
  • Lewis Hammond
  • Eric Bigelow
  • Alexander Pan
  • Lauro Langosco
  • Tomasz Korbak
  • Heidi Zhang
  • Ruiqi Zhong
  • Seán Ó hÉigeartaigh
  • Gabriel Recchia
  • Giulio Corsi
  • Alan Chan
  • Markus Anderljung
  • Lilian Edwards
  • Aleksandar Petrov
  • Christian Schroeder de Witt
  • Sumeet Ramesh Motwan
  • Yoshua Bengio
  • Danqi Chen
  • Philip H. S. Torr
  • Samuel Albanie
  • Tegan Maharaj
  • Jakob Foerster
  • Florian Tramer
  • He He
  • Atoosa Kasirzadeh
  • Yejin Choi
  • David Krueger

Abstract

  • This work identifies 18 foundational challenges in assuring the alignment and safety of large language models (LLMs). These challenges are organized into three different categories: scientific understanding of LLMs, development and deployment methods, and sociotechnical challenges. Based on the identified challenges, we pose $200+$ concrete research questions.

Translation (by gpt-4o-mini)

  • 本研究では、大規模言語モデル(LLMs)の整合性と安全性を確保する上での18の基盤的な課題を特定しました。これらの課題は、LLMsの科学的理解、開発および展開方法、社会技術的課題の3つの異なるカテゴリに整理されています。特定された課題に基づき、200以上の具体的な研究質問を提起します。

Summary (by gpt-4o-mini)

  • 本研究では、LLMsの整合性と安全性に関する18の基盤的課題を特定し、科学的理解、開発・展開方法、社会技術的課題の3つのカテゴリに整理。これに基づき、200以上の具体的な研究質問を提起。

Metadata

Metadata

Assignees

No one assigned

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions