On the Sensitivity of Reward Inference to Misspecified Human Models

Hong, Joey; Bhatia, Kush; Dragan, Anca

Computer Science > Machine Learning

arXiv:2212.04717 (cs)

[Submitted on 9 Dec 2022 (v1), last revised 30 Oct 2023 (this version, v2)]

Title:On the Sensitivity of Reward Inference to Misspecified Human Models

Authors:Joey Hong, Kush Bhatia, Anca Dragan

View PDF

Abstract:Inferring reward functions from human behavior is at the center of value alignment - aligning AI objectives with what we, humans, actually want. But doing so relies on models of how humans behave given their objectives. After decades of research in cognitive science, neuroscience, and behavioral economics, obtaining accurate human models remains an open research topic. This begs the question: how accurate do these models need to be in order for the reward inference to be accurate? On the one hand, if small errors in the model can lead to catastrophic error in inference, the entire framework of reward learning seems ill-fated, as we will never have perfect models of human behavior. On the other hand, if as our models improve, we can have a guarantee that reward accuracy also improves, this would show the benefit of more work on the modeling side. We study this question both theoretically and empirically. We do show that it is unfortunately possible to construct small adversarial biases in behavior that lead to arbitrarily large errors in the inferred reward. However, and arguably more importantly, we are also able to identify reasonable assumptions under which the reward inference error can be bounded linearly in the error in the human model. Finally, we verify our theoretical insights in discrete and continuous control tasks with simulated and human data.

Comments:	published as a paper in ICLR 2023; 17 pages, 12 figures
Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2212.04717 [cs.LG]
	(or arXiv:2212.04717v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2212.04717

Submission history

From: Joey Hong [view email]
[v1] Fri, 9 Dec 2022 08:16:20 UTC (8,326 KB)
[v2] Mon, 30 Oct 2023 05:01:12 UTC (8,321 KB)

Computer Science > Machine Learning

Title:On the Sensitivity of Reward Inference to Misspecified Human Models

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:On the Sensitivity of Reward Inference to Misspecified Human Models

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators