This repository supports the paper titled "Towards Bidirectional Human-AI Alignment: A Systematic Review for Clarifications, Framework, and Future Directions".
In this paper, we proposed the "Bidirectional Human-AI Alignment" framework. The overview figure of framework provided below includes a clarified definition based on a systematic review of over 400 papers spanning the domains of Human-Computer Interaction (HCI), Natural Language Processing (NLP), and Machine Learng (ML), and more.
Definition. Bidirectional Human-AI Alignment is a comprehensive framework that encompasses two interconnected alignment processes: (A) Aligning AI to Humans and (B) Aligning Humans to AI. The former focuses on integrating human specifications to train, steer, and customize AI, while the latter investigates human cognitive and behavioral adaptations to AI, which supports humans in understanding, critiquing, collaborating with, and adapting to AI advancements.
Contributor: Hua Shen. Please see the Paper and a full list of people at Link.
- 1. Alignment Basics
- 2. Human Values and Specifications
- 3. Integrating Human Specifications into AI
- 4. Human Cognitive Adjustment to AI
- 5. Human Adaptive Behavior to AI
- People and Community
- Contributions
- How to Cite the Repo
The overview papers are organized in chronological order.
-
Open problems and fundamental limitations of reinforcement learning from human feedback. Casper, Stephen, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman et al., Sep 2023. [pdf]
-
AI alignment in the design of interactive AI: Specification alignment, process alignment, and evaluation support. Terry, Michael, Chinmay Kulkarni, Martin Wattenberg, Lucas Dixon, and Meredith Ringel Morris., Oct 2023. [pdf]
-
A roadmap to pluralistic alignment. Sorensen, Taylor, Jared Moore, Jillian Fisher, Mitchell Gordon, Niloofar Mireshghallah, Christopher Michael Rytting, Andre Ye et al., Feb 2024. [pdf]
-
Foundational challenges in assuring alignment and safety of large language models. Anwar, Usman, Abulhair Saparov, Javier Rando, Daniel Paleka, Miles Turpin, Peter Hase, Ekdeep Singh Lubana et al., Apr 2024. [pdf]
-
Managing extreme AI risks amid rapid progress. Bengio, Yoshua, Geoffrey Hinton, Andrew Yao, Dawn Song, Pieter Abbeel, Trevor Darrell, Yuval Noah Harari et al., May 2024. [pdf]
-
Towards Bidirectional Human-AI Alignment: A Systematic Review for Clarifications, Framework, and Future Directions. Shen, Hua, Tiffany Knearem, Reshmi Ghosh, Kenan Alkiek, Kundan Krishna, Yachuan Liu, Ziqiao Ma et al., Jun 2024. [pdf]
-
Artificial intelligence, values, and alignment. Gabriel, Iason., 2020. [pdf]
-
(Schwartz Theory of Basic Values) Are there universal aspects in the structure and contents of human values?. Schwartz, Shalom H., 1994. [pdf]
-
(Schwartz Theory of Basic Values) An overview of the Schwartz theory of basic values. Schwartz, Shalom H., 2012. [pdf]
-
(Moral Foundation Theory) Moral foundations theory: The pragmatic validity of moral pluralism. Graham, Jesse, Jonathan Haidt, Sena Koleva, Matt Motyl, Ravi Iyer, Sean P. Wojcik, and Peter H. Ditto., 2013. [pdf]
Individuals:
-
(2019 CHI) Improving fairness in machine learning systems: What do industry practitioners need?. Holstein, Kenneth, Jennifer Wortman Vaughan, Hal Daumé III, Miro Dudik, and Hanna Wallach. [pdf]
-
(2022 Arxiv) Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned. Ganguli, Deep, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, Saurav Kadavath, Ben Mann et al. arXiv preprint arXiv:2209.07858 (2022).
-
(2023 EMNLP) Ethical reasoning over moral alignment: A case and framework for in-context ethical policies in LLMs. Rao, Abhinav, Aditi Khandelwal, Kumar Tanmay, Utkarsh Agarwal, and Monojit Choudhury. [pdf]
-
(2021 CHI) Human perceptions on moral responsibility of AI: A case study in AI-assisted bail decision-making. Lima, Gabriel, Nina Grgić-Hlača, and Meeyoung Cha. [pdf]
Society:
-
(2023 EMNLP) Do llms understand social knowledge? evaluating the sociability of large language models with socket benchmark. Choi, Minje, Jiaxin Pei, Sagar Kumar, Chang Shu, and David Jurgens. [pdf]
-
(2023 CHI) Competent but Rigid: Identifying the Gap in Empowering AI to Participate Equally in Group Decision-Making. Zheng, Chengbo, Yuheng Wu, Chuhan Shi, Shuai Ma, Jiehui Luo, and Xiaojuan Ma.
-
(2023 ACL) The ecological fallacy in annotation: Modelling human label variation goes beyond sociodemographics. Orlikowski, Matthias, Paul Röttger, Philipp Cimiano, and Dirk Hovy. [pdf]
-
(2023 Arxiv) TASRA: a taxonomy and analysis of societal-scale risks from AI. Critch, Andrew, and Stuart Russell. arXiv preprint arXiv:2306.06924 (2023). [pdf]
-
(2020 EMNLP) Social chemistry 101: Learning to reason about social and moral norms. Forbes, Maxwell, Jena D. Hwang, Vered Shwartz, Maarten Sap, and Yejin Choi. [pdf]
Interaction:
-
(2023 EMNLP) FANToM: A benchmark for stress-testing machine theory of mind in interactions. Kim, Hyunwoo, Melanie Sclar, Xuhui Zhou, Ronan Le Bras, Gunhee Kim, Yejin Choi, and Maarten Sap. [pdf]
-
(2024 NeurIPS) Lima: Less is more for alignment. Zhou, Chunting, Pengfei Liu, Puxin Xu, Srinivasan Iyer, Jiao Sun, Yuning Mao, Xuezhe Ma et al. [pdf]
Self-Enhancement:
-
(2023 CSCW) Deliberating with AI: improving decision-making for the future through participatory AI design and stakeholder deliberation. Zhang, Angie, Olympia Walker, Kaci Nguyen, Jiajun Dai, Anqing Chen, and Min Kyung Lee. [pdf]
-
(2023 CHI) “Should I Follow the Human, or Follow the Robot?”—Robots in Power Can Have More Influence Than Humans on Decision-Making. Hou, Yoyo Tsung-Yu, Wen-Ying Lee, and Malte Jung. [pdf]
Self-Transcendence:
-
(2023 NeurIPS) Auditing for human expertise. Alur, Rohan, Loren Laine, Darrick Li, Manish Raghavan, Devavrat Shah, and Dennis Shung. [pdf]
-
(2022 Arxiv) Training a helpful and harmless assistant with reinforcement learning from human feedback. Bai, Yuntao, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain et al. [pdf]
Conservation:
-
(2022 Arxiv) Constitutional ai: Harmlessness from ai feedback. Bai, Yuntao, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen et al. [pdf]
-
(2024, ICLR) Fine-tuning aligned language models compromises safety, even when users do not intend to!. Qi, Xiangyu, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. [pdf]
Openness to Change:
-
(2023 ACL) Increasing diversity while maintaining accuracy: Text data generation with large language models and human interventions. Chung, John Joon Young, Ece Kamar, and Saleema Amershi. [pdf]
-
(2024 ICLR) Fine-tuning language models for factuality. Tian, Katherine, Eric Mitchell, Huaxiu Yao, Christopher D. Manning, and Chelsea Finn. [pdf]
Desired Values from AI Tools:
-
(2024 ICLR) Generative judge for evaluating alignment. Li, Junlong, Shichao Sun, Weizhe Yuan, Run-Ze Fan, Hai Zhao, and Pengfei Liu. [pdf]
-
(2022 ACL) Are shortest rationales the best explanations for human understanding?. Shen, Hua, Tongshuang Wu, Wenbo Guo, and Ting-Hao'Kenneth Huang. [pdf]
-
(2022 NeurIPS) Training language models to follow instructions with human feedback. Ouyang, Long, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang et al. [pdf]
-
(2023 ACL) Rl4f: Generating natural language feedback with reinforcement learning for repairing model outputs. Akyürek, Afra Feyza, Ekin Akyürek, Aman Madaan, Ashwin Kalyan, Peter Clark, Derry Wijaya, and Niket Tandon. [pdf]
-
(2023 IUI) Constitutionmaker: Interactively critiquing large language models by converting feedback into principles. Petridis, Savvas, Benjamin D. Wedin, James Wexler, Mahima Pushkarna, Aaron Donsbach, Nitesh Goyal, Carrie J. Cai, and Michael Terry. [pdf]
-
(2023 EMNLP) Towards a holistic landscape of situated theory of mind in large language models. Ma, Ziqiao, Jacob Sansom, Run Peng, and Joyce Chai. [pdf]
-
(2023 EMNLP) Hi-tom: A benchmark for evaluating higher-order theory of mind reasoning in large language models. He, Yinghui, Yufan Wu, Yilin Jia, Rada Mihalcea, Yulong Chen, and Naihao Deng. [pdf]
-
(2021 EMNLP) A scalable framework for learning from implicit user feedback to improve natural language understanding in large-scale conversational ai systems. Park, Sunghyun, Han Li, Ameen Patel, Sidharth Mudgal, Sungjin Lee, Young-Bum Kim, Spyros Matsoukas, and Ruhi Sarikaya. [pdf]
-
(2024 ICLR) Identifying the risks of lm agents with an lm-emulated sandbox. Ruan, Yangjun, Honghua Dong, Andrew Wang, Silviu Pitis, Yongchao Zhou, Jimmy Ba, Yann Dubois, Chris J. Maddison, and Tatsunori Hashimoto. [pdf]
-
(2023 EMNLP) Aligning large language models through synthetic feedback. Kim, Sungdong, Sanghwan Bae, Jamin Shin, Soyoung Kang, Donghyun Kwak, Kang Min Yoo, and Minjoon Seo. [pdf]
-
(2023 EMNLP) Multiturncleanup: A benchmark for multi-turn spoken conversational transcript cleanup., Shen, Hua, Vicky Zayats, Johann C. Rocholl, Daniel D. Walker, and Dirk Padfield. [pdf]
-
(2023 EMNLP) Coannotating: Uncertainty-guided work allocation between human and large language models for data annotation. Li, Minzhi, Taiwei Shi, Caleb Ziems, Min-Yen Kan, Nancy F. Chen, Zhengyuan Liu, and Diyi Yang. [pdf]
-
(2024 NeurIPS) RRHF: Rank responses to align language models with human feedback. Yuan, Hongyi, Zheng Yuan, Chuanqi Tan, Wei Wang, Songfang Huang, and Fei Huang. [pdf]
-
(2024 NeurIPS) Direct preference optimization: Your language model is secretly a reward model. Rafailov, Rafael, Archit Sharma, Eric Mitchell, Christopher D. Manning, Stefano Ermon, and Chelsea Finn. [pdf]
-
(2023 TMLR) Raft: Reward ranked finetuning for generative foundation model alignment. Dong, Hanze, Wei Xiong, Deepanshu Goyal, Yihan Zhang, Winnie Chow, Rui Pan, Shizhe Diao, Jipeng Zhang, Kashun Shum, and Tong Zhang. [pdf]
-
(2023 EMNLP) Humanoid agents: Platform for simulating human-like generative agents. Wang, Zhilin, Yu Ying Chiu, and Yu Cheung Chiu. [pdf]
-
(2023 ICLR) Critic: Large language models can self-correct with tool-interactive critiquing. Gou, Zhibin, Zhihong Shao, Yeyun Gong, Yelong Shen, Yujiu Yang, Nan Duan, and Weizhu Chen. [pdf]
-
(2024 ICLR) Quality-diversity through ai feedback. Bradley, Herbie, Andrew Dai, Hannah Teufel, Jenny Zhang, Koen Oostermeijer, Marco Bellagente, Jeff Clune, Kenneth Stanley, Grégory Schott, and Joel Lehman. [pdf]
- (2023 ACL) Winoqueer: A community-in-the-loop benchmark for anti-lgbtq+ bias in large language models. Felkner, Virginia K., Ho-Chun Herbert Chang, Eugene Jang, and Jonathan May. [pdf]
-
(2024 ICLR) Group preference optimization: Few-shot alignment of large language models. Zhao, Siyan, John Dang, and Aditya Grover. [pdf]
-
(2022 AAMAS) MORAL: Aligning AI with human norms through multi-objective reinforced active learning. Peschl, Markus, Arkady Zgonnikov, Frans A. Oliehoek, and Luciano C. Siebert. [pdf]
-
(2022 EMNLP) BERTScore is unfair: On social bias in language model-based metrics for text generation. Sun, Tianxiang, Junliang He, Xipeng Qiu, and Xuanjing Huang. [pdf]
-
(2022 ICASSP) Improving fairness in speaker verification via group-adapted fusion network. Shen, Hua, Yuguang Yang, Guoli Sun, Ryan Langman, Eunjung Han, Jasha Droppo, and Andreas Stolcke. [pdf]
-
(2023 EMNLP) LDM
$^ 2$ : A Large Decision Model Imitating Human Cognition with Dynamic Memory Enhancement. Wang, Xingjin, Linjing Li, and Daniel Zeng. [pdf]
-
(2023 ICLR) Personalized reward learning with interaction-grounded learning (IGL). Maghakian, Jessica, Paul Mineiro, Kishan Panaganti, Mark Rucker, Akanksha Saran, and Cheng Tan. [pdf]
-
(2023 EMNLP) Improving diversity of demographic representation in large language models via collective-critiques and self-voting. Lahoti, Preethi, Nicholas Blumm, Xiao Ma, Raghavendra Kotikalapudi, Sahitya Potluri, Qijun Tan, Hansa Srinivasan et al. [pdf]
-
(2022 ACL) Leveraging similar users for personalized language modeling with limited data. Welch, Charles, Chenxi Gu, Jonathan K. Kummerfeld, Verónica Pérez-Rosas, and Rada Mihalcea. [pdf]
-
(2022 NAACL) Aligning generative language models with human values. Liu, Ruibo, Ge Zhang, Xinyu Feng, and Soroush Vosoughi. [pdf]
-
(2023 ACL) Finspector: A human-centered visual inspection tool for exploring and comparing biases among foundation models. Kwon, Bum Chul, and Nandana Mihindukulasooriya. [pdf]
-
(2024 NeurIPS) Principle-driven self-alignment of language models from scratch with minimal human supervision. Sun, Zhiqing, Yikang Shen, Qinhong Zhou, Hongxin Zhang, Zhenfang Chen, David Cox, Yiming Yang, and Chuang Gan. [pdf]
-
(2023 EMNLP) Can Large Language Models Capture Dissenting Human Voices?. Lee, Noah, Na Min An, and James Thorne. [pdf]
-
(2023 NeurIPS) Alpacafarm: A simulation framework for methods that learn from human feedback. Dubois, Yann, Chen Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy S. Liang, and Tatsunori B. Hashimoto. [pdf]
-
(2023 EMNLP) Gentopia: A collaborative platform for tool-augmented llms. Xu, Binfeng, Xukun Liu, Hua Shen, Zeyu Han, Yuhan Li, Murong Yue, Zhiyuan Peng, Yuchen Liu, Ziyu Yao, and Dongkuan Xu. [pdf]
-
(2022 EMNLP) Potato: The portable text annotation tool. Pei, Jiaxin, Aparna Ananthasubramaniam, Xingyao Wang, Naitian Zhou, Jackson Sargent, Apostolos Dedeloudis, and David Jurgens. [pdf]
-
(2024 ICLR) Uni-RLHF: Universal Platform and Benchmark Suite for Reinforcement Learning with Diverse Human Feedback. Yuan, Yifu, Jianye Hao, Yi Ma, Zibin Dong, Hebin Liang, Jinyi Liu, Zhixin Feng, Kai Zhao, and Yan Zheng. [pdf]
-
(2022 CHI) What is AI literacy? Competencies and design considerations. Long, Duri, and Brian Magerko. [pdf]
-
(2020 CSCW) Intersectional AI: A study of how information science students think about ethics and their impact. McDonald, Nora, and Shimei Pan. [pdf]
-
(2023 CHI) Anglekindling: Supporting journalistic angle ideation with large language models. Petridis, Savvas, Nicholas Diakopoulos, Kevin Crowston, Mark Hansen, Keren Henderson, Stan Jastrzebski, Jeffrey V. Nickerson, and Lydia B. Chilton. [pdf]
-
(2020 EMNLP) The language interpretability tool: Extensible, interactive visualizations and analysis for NLP models. Tenney, Ian, James Wexler, Jasmijn Bastings, Tolga Bolukbasi, Andy Coenen, Sebastian Gehrmann, Ellen Jiang et al. [pdf]
-
(2023 CSCW Demo) ConvXAI: Delivering heterogeneous AI explanations via conversations to support human-AI scientific writing. Shen, Hua, Chieh-Yang Huang, Tongshuang Wu, and Ting-Hao Kenneth Huang. [pdf]
-
(2019 CHI) Understanding the effect of accuracy on trust in machine learning models. Yin, Ming, Jennifer Wortman Vaughan, and Hanna Wallach. [pdf]
-
(2023 IUI) Appropriate reliance on AI advice: Conceptualization and the effect of explanations. Schemmer, Max, Niklas Kuehl, Carina Benz, Andrea Bartos, and Gerhard Satzger. [pdf]
-
(2024 CHI) An AI-Resilient Text Rendering Technique for Reading and Skimming Documents. Gu, Ziwei, Ian Arawjo, Kenneth Li, Jonathan K. Kummerfeld, and Elena L. Glassman. [pdf]
-
(2020 CHI) Co-designing checklists to understand organizational challenges and opportunities around fairness in AI. Madaio, Michael A., Luke Stark, Jennifer Wortman Vaughan, and Hanna Wallach. [pdf]
-
(2021 CSCW) Problematic machine behavior: A systematic literature review of algorithm audits. Bandy, Jack. [pdf]
-
(2019 CHI) Will you accept an imperfect ai? exploring designs for adjusting end-user expectations of ai systems. Kocielnik, Rafal, Saleema Amershi, and Paul N. Bennett. [pdf]
-
(2023 CHI) Measuring and understanding trust calibrations for automated systems: a survey of the state-of-the-art and future directions. Wischnewski, Magdalena, Nicole Krämer, and Emmanuel Müller. [pdf]
-
(2022 CHI) Ai chains: Transparent and controllable human-ai interaction by chaining large language model prompts. Wu, Tongshuang, Michael Terry, and Carrie Jun Cai. [pdf]
-
(2023 CHI) Patat: Human-ai collaborative qualitative coding with explainable interactive rule synthesis. Gebreegziabher, Simret Araya, Zheng Zhang, Xiaohang Tang, Yihao Meng, Elena L. Glassman, and Toby Jia-Jun Li. [pdf]
-
(2019 CHI) Aila: Attentive interactive labeling assistant for document classification through attention-based deep neural networks. Choi, Minsuk, Cheonbok Park, Soyoung Yang, Yonggyu Kim, Jaegul Choo, and Sungsoo Ray Hong. [pdf]
-
(2023 UIST) Generative agents: Interactive simulacra of human behavior. Park, Joon Sung, Joseph O'Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. [pdf]
-
(2021 CSCW) The design of reciprocal learning between human and artificial intelligence. Zagalsky, Alexey, Dov Te'eni, Inbal Yahav, David G. Schwartz, Gahl Silverman, Daniel Cohen, Yossi Mann, and Dafna Lewinsky. [pdf]
-
(2023 CHI) AI knowledge: Improving AI delegation through human enablement. Pinski, Marc, Martin Adam, and Alexander Benlian. [pdf]
-
(2022 HCR) Rethinking communication in the era of artificial intelligence. Sundar, S. Shyam, and Eun-Ju Lee. [pdf]
-
(2022 CHI) How experienced designers of enterprise applications engage AI as a design material. Yildirim, Nur, Alex Kass, Teresa Tung, Connor Upton, Donnacha Costello, Robert Giusti, Sinem Lacin et al. [pdf]
-
(2024 AIED) How to Teach Programming in the AI Era? Using LLMs as a Teachable Agent for Debugging. Ma, Qianou, Hua Shen, Kenneth Koedinger, and Sherry Tongshuang Wu. [pdf]
-
(2024 CHI) Rehearsal: Simulating conflict to teach conflict resolution. Shaikh, Omar, Valentino Emil Chai, Michele Gelfand, Diyi Yang, and Michael S. Bernstein. [pdf]
-
(2020 HCOMP) How useful are the machine-generated interpretations to general users? a human evaluation on guessing the incorrectly predicted labels. Shen, Hua, and Ting-Hao Huang. [pdf]
-
(2024 Arxiv) How AI Ideas Affect the Creativity, Diversity, and Evolution of Human Ideas: Evidence From a Large, Dynamic Experiment. Ashkinaze, Joshua, Julia Mendelsohn, Li Qiwei, Ceren Budak, and Eric Gilbert. [pdf]
-
(2022 IUI) Do people engage cognitively with AI? Impact of AI assistance on incidental learning. Gajos, Krzysztof Z., and Lena Mamykina. [pdf]
-
(2023 CSCW) Remove, reduce, inform: what actions do people want Social Media platforms to take on potentially misleading content?. Atreja, Shubham, Libby Hemphill, and Paul Resnick. [pdf]
-
(2023 ICLR) Towards interpretable deep reinforcement learning with human-friendly prototypes. Kenny, Eoin M., Mycal Tucker, and Julie Shah. [pdf]
-
(2022 Information Processing & Management) Human vs. AI: Understanding the impact of anthropomorphism on consumer response to chatbots from the perspective of trust and relationship norms. Cheng, Xusen, Xiaoping Zhang, Jason Cohen, and Jian Mou. [pdf]
-
(2021 CHI) Human-AI interaction in human resource management: Understanding why employees resist algorithmic evaluation at workplaces and how to mitigate burdens. Park, Hyanghee, Daehwan Ahn, Kartik Hosanagar, and Joonhwan Lee. [pdf]
-
(2023 FAccT) Regulating ChatGPT and other large generative AI models. Hacker, Philipp, Andreas Engel, and Marco Mauer. [pdf]
-
(2023 FAccT) Who Should Pay When Machines Cause Harm? Laypeople’s Expectations of Legal Damages for Machine-Caused Harm. Lima, Gabriel, Nina Grgic-Hlaca, Jin Keun Jeong, and Meeyoung Cha. [pdf]
-
(2021 CHI) Human-AI interaction in human resource management: Understanding why employees resist algorithmic evaluation at workplaces and how to mitigate burdens. Park, Hyanghee, Daehwan Ahn, Kartik Hosanagar, and Joonhwan Lee. [pdf]
-
(2020 CHI) Co-designing checklists to understand organizational challenges and opportunities around fairness in AI. Madaio, Michael A., Luke Stark, Jennifer Wortman Vaughan, and Hanna Wallach. [pdf]
-
(2024 Arxiv) The RealHumanEval: Evaluating Large Language Models' Abilities to Support Programmers. Mozannar, Hussein, Valerie Chen, Mohammed Alsobay, Subhro Das, Sebastian Zhao, Dennis Wei, Manish Nagireddy, Prasanna Sattigeri, Ameet Talwalkar, and David Sontag. [pdf]
-
(2023 EMNLP) Harnessing the power of LLMs: Evaluating human-AI text co-creation through the lens of news headline generation. Ding, Zijian, Alison Smith-Renner, Wenjuan Zhang, Joel R. Tetreault, and Alejandro Jaimes. [pdf]
-
(2023 CCS) Do users write more insecure code with AI assistants? Neil Perry, Megha Srivastava, Deepak Kumar, and Dan Boneh. [pdf]
-
(2023 CSCW) Deliberating with AI: improving decision-making for the future through participatory AI design and stakeholder deliberation. Zhang, Angie, Olympia Walker, Kaci Nguyen, Jiajun Dai, Anqing Chen, and Min Kyung Lee. [pdf]
-
(2021 CHI) Towards mutual theory of mind in human-ai interaction: How language reflects what students perceive about a virtual teaching assistant. Wang, Qiaosi, Koustuv Saha, Eric Gregori, David Joyner, and Ashok Goel. [pdf]
-
(2023 ICML) Whose opinions do language models reflect?. Santurkar, Shibani, Esin Durmus, Faisal Ladhak, Cinoo Lee, Percy Liang, and Tatsunori Hashimoto. [pdf]
-
(2024 PNAS) Generative artificial intelligence, human creativity, and art. Zhou, Eric, and Dokyun Lee. [pdf]
- Tiffany Knearem (Google) [link]
- Reshmi Ghosh (Microsoft) [link]
- Kenan Alkiek (UMich) [link]
- Kundan Krishna (CMU) [link]
- Yachuan Liu (UMich) [link]
- Ziqiao (Martin) Ma (UMich) [link]
- Savvas Petridis (Google DeepMind) [link]
- Yi-Hao Peng (CMU) [link]
- Li Qiwei (UMich) [link]
- Sushrita Rakshit (UMich) [link]
- Chenglei Si (Stanford) [link]
- Yutong Xie (UMich) [link]
- Jeffrey P. Bigham (CMU) [link]
- Frank Bentley (Google) [link]
- Joyce Chai (UMich) [link]
- Zachary Lipton (CMU) [link]
- Qiaozhu Mei (UMich) [link]
- Rada Mihalcea (UMich) [link]
- Michael Terry (Google DeepMind) [link]
- Diyi Yang (Stanford) [link]
- Meredith Ringel Morris (Google DeepMind) [link]
- Paul Resnick (UMich) [link]
- David Jurgens (UMich) [link]
- We welcome more people who are interested in contributing to bidirectional human-AI alignment to join our community!
There are many more papers on Bidirectional Human-AI Alignment beyond the ones listed above. Due to limitations, we cannot cover all relevant papers. If you know of any highly relevant papers that should be included, please contact Hua Shen. Contributions to both Bidirectional Alignment and this reading list are welcome. Thank you!
If you find the repository helpful to your research and would like to cite it, please see the bibtex below:
@article{shen2024towards,
title={Towards Bidirectional Human-AI Alignment: A Systematic Review for Clarifications, Framework, and Future Directions},
author={Shen, Hua and Knearem, Tiffany and Ghosh, Reshmi and Alkiek, Kenan and Krishna, Kundan and Liu, Yachuan and Ma, Ziqiao and Petridis, Savvas and Peng, Yi-Hao and Qiwei, Li and Rakshit, Sushrita and Si, Chenglei and Xie, Yutong and Bigham, Jeffrey P. and Bentley, Frank and Chai, Joyce and Lipton, Zachary and Mei, Qiaozhu and Mihalcea, Rada and Terry, Michael and Yang, Diyi and Morris, Meredith Ringel and Resnick, Paul and Jurgens, David},
journal={arXiv preprint arXiv:2406.09264},
year={2024}
}