paper · First submission: March 9, 2019
Literal or Pedagogic Human? Analyzing Human Model Misspecification in Objective Learning
Teaching assumptions can improve modeled cooperation while making reward inference brittle when real people behave differently.
- Revision: June 29, 2019
- Proceedings: 2020
Smitha Milli and Anca D. Dragan examine whether a learner should assume a person is performing a task or teaching it. This entry reviews the PMLR proceedings original for UAI 2019, published in the 2020 volume; the chronology marks the first arXiv submission.[1]
The mismatch
In Figure 1, walking on pavement might simply be a convenient route. A learner assuming deliberate teaching can instead infer that grass is forbidden. A literal learner treats behavior as task performance; a pedagogic learner assumes the person chose actions to inform it. Neither label describes obedience to commands.[1]
In a common-payoff game, Claim 2.1 ranks four pairings under recursive robot best responses and human improving responses. A pedagogic learner paired with a literal human has no higher payoff than a literal learner paired with a pedagogic human. The correctly matched pedagogic pair can still do best: this is an asymmetry of mismatch, not universal superiority of literal inference (§2).[1]
Evidence from human demonstrations
The authors reanalyze gridworld demonstrations from an earlier experiment with 60 Mechanical Turk participants. Reward-inference accuracy is worse when a pedagogic learner misreads literal behavior. Even with people instructed to teach, the literal learner slightly outperforms the pedagogic one, but that difference is not statistically significant (§3; Figure 3). Mixture models partly explain the discrepancy; improved behavior prediction need not improve inference of the hidden reward (§§4–5).[1]
Historical context
The paper builds on Cooperative Inverse Reinforcement Learning. It sharpens Human Model Misspecification: learning a goal shifts some specification work into the interpretation of people. Its human-data result does not establish robustness across other tasks, cultures or deployed assistants.[1]
Sources
Pages that link here
Last updated 2026-10-10