Interpretability

2026

Isaac Song, Mohammed Rehan Parwani, Glenn Matlin, Emile Timothy Anand, Akhil Theerthala, Arjun Chatterjee, Anthony Wen-Ming Zang, Maria Kostylew, Yonadav G. Shavit, Sebastian Krier, and Mark O. Riedl
Role Steering of Language Models for Social Simulations
Proceedings of the COLM 2026 Workshop on Social Simulations with LLMs (2026).
arXiv Workshop bibtexSocial SimulationAgentsLarge Language ModelsInterpretability

Glenn Matlin, Chandreyi Chakraborty, Saehee Eom, Mika Okamoto, Rayan Castilla, Louis Jaburi, Alvin Deng, Taywon Min, Lucia Quirke, Stella Biderman, and Mark Riedl
Capability Provenance in Language Models: A Case Study in Social Reasoning
Proceedings of the 2026 Conference on Language Models (2026).
arXiv Conference bibtexLarge Language ModelsInterpretability

2024

Kenneth Eaton, Jonathan Balloch, Julia Kim, and Mark Riedl
The Interpretability of Codebooks in Model-Based Reinforcement Learning is Limited
Proceedings of the 2024 Reinforcement Learning Conference Workshop I Can't Believe It's not Better (2024).
arXiv Workshop bibtexAgentsExplainable AIInterpretabilityReinforcement Learning

2022

Xiangyu Peng, Mark O. Riedl, and Prithviraj Ammanabrolu
Inherently Explainable Reinforcement Learning in Natural Language
Proceedings of NeurIPS 2022 (2022).
arXiv Conference bibtexInteractive StoriesAgentsExplainable AIInterpretabilityReinforcement Learning

Xiangyu Peng, Mark O. Riedl, and Prithviraj Ammanabrolu
Inherently Explainable Reinforcement Learning in Natural Language
Proceedings of the 2022 Multi-disciplinary Conference on Reinforcement Learning and Decision Making (2022).
arXiv Conference bibtexInteractive StoriesAgentsExplainable AIInterpretabilityReinforcement Learning

2021

Xiangyu Peng, Prithviraj Ammanabrolu, and Mark Riedl
Explainable Reinforcement Learning Agents with Stacked Hierarchical Graph Attention
Workshop on Explainable Graph-based Machine Learning at AKBC (2021).
Workshop bibtexAgentsExplainable AIInterpretabilityReinforcement Learning

Sarah Wiegreffe, Ana Marasovic, and Noah A. Smith
Measuring Association Between Labels and Free-Text Rationales
Proceedings of NAACL 2021 (2021).
arXiv Conference bibtexLarge Language ModelsExplainable AIInterpretability