AI Safety

2026

Glenn Matlin, Isaac Song, Anthony Wen-Ming Zang, and Mark O. Riedl
No One Wins in Nuclear War: Social Simulations of High-Stakes Military Decision-Making
Proceedings of the COLM 2026 Workshop on Social Simulations with LLMs (2026).
arXiv Workshop bibtexSocial SimulationAgentsLarge Language ModelsAI Safety

Mark O. Riedl and Glenn Matlin
Position: AI is Not Ready for Strategic Conflict
Proceedings of the COLM 2026 Workshop on Social Simulations with LLMs (2026).
Workshop bibtexSocial SimulationAgentsPolicy and LawAI Safety

Geigh Zollicoffer, Tanush Chopra, Mingkuan Yan, Xiaoxu Ma, Kenneth Eaton, and Mark Riedl
World Model Robustness via Surprise Recognition
Proceedings of the 2026 CVPR Findings (2026).
arXiv Conference bibtexAgentsReinforcement LearningAI Safety

2025

Glenn Matlin, Parv Mahajan, Isaac Song, Yixiong Hao, Ryan Bard, Stu Topp, Evan Montoya, M. Rehan Parwani, Soham Shetty, and Mark Riedl
Shall We Play a Game? Language Models for Open-ended Wargames
Proceedings of the 2025 EMNLP Word Play Workshop (2025).
arXiv Workshop bibtexSocial SimulationAgentsAI Safety

Geigh Zollicoffer, Kenneth Eaton, Jonathan Balloch, Julia Kim, Riedl Mark O., and Robert Wright
Novelty Detection in Reinforcement Learning with World Models
International Conference on Machine Learning (2025).
arXiv Conference bibtexAgentsReinforcement LearningAI Safety

2024

Ashutosh Baheti, Ximing Lu, Faeze Brahman, Ronan Le Bras, Maarten Sap, and Mark O. Riedl
Improving Language Models with Advantage-based Offline Policy Gradients
Proceedings of ICLR 2024 (2024).
arXiv Conference bibtexLarge Language ModelsReinforcement LearningValue AlignmentAI Safety

2021

Ashutosh Baheti, Maarten Sap, Alan Ritter, and Mark O. Riedl
Just Say No: Analyzing the Stance of Neural Dialogue Generation in Offensive Contexts
Proceedings of EMNLP 2021 (2021).
arXiv Conference bibtexLarge Language ModelsValue AlignmentAI Safety

Md Sultan Al Nahian, Spencer Frazier, Brent Harrison, and Mark O. Riedl
Training Value-Aligned Reinforcement Learning Agents Using a Normative Prior
arXiv:2104.09469 (2021).
arXiv bibtexAgentsReinforcement LearningValue AlignmentAI Safety

2020

Xiangyu Peng, S. Li, Spencer Frazier, and Mark O. Riedl
Reducing Non-Normative Text Generation from Language Models
International Conference on Natural Language Generation (2020).
arXiv Conference bibtexLarge Language ModelsValue AlignmentAI Safety

2017

Mark O. Riedl and Brent. Harrison
Enter the Matrix: A Virtual World Approach to Safely Interruptable Autonomous Systems
Proceedings of the AAAI 2017 Workshop on SafeAI (2017).
arXiv Workshop bibtexAgentsReinforcement LearningAI Safety

2016

Mark Riedl and Brent Harrison
Using Stories to Teach Human Values to Artificial Agents
Proceedings of the 2nd International Workshop on AI, Ethics and Society (2016).
PDF Workshop bibtexAgentsValue AlignmentAI Safety

Brent Harrsion and Mark O Riedl
Learning From Stories: Using Crowdsourced Narratives to Train Virtual Agents
Proceedings of the 2016 AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment (2016).
PDF Conference bibtexAgentsValue AlignmentAI SafetyHuman Computation

Brent Harrison and Mark Riedl
Towards Learning From Stories: An Approach for Interactive Machine Learning
Proceedings of the AAAI'16 Workshop on Symbiotic Cognitive Systems (2016).
PDF Workshop bibtexAgentsValue AlignmentAI Safety

Brent Harrison, Siddhartha Banerjee, and Mark O. Riedl
Learning from Stories: Using Natural Communication to Train Believable Agents
Proceedings of the 2016 IJCAI Workshop on Interactive Machine Learning (2016).
PDF Workshop bibtexAgentsValue AlignmentAI Safety