concept
Human Model Misspecification
Errors in what an assistant assumes people value, know or intend when interpreting their behavior.
Learning from a person requires an interpretation of that person’s behavior. A system might assume task competence, deliberate teaching, a particular noise level or a fixed family of preferences. Human model misspecification occurs when such assumptions do not fit the person or interaction. Uncertainty among the model’s possibilities does not ensure that the right possibility is represented.[3][4]
Performing, teaching and commanding
Cooperative Inverse Reinforcement Learning makes teaching part of cooperative assistance.[5] Literal or Pedagogic Human? Analyzing Human Model Misspecification in Objective Learning shows why assuming teaching can also be brittle. Its pavement example distinguishes an ordinary route choice from a purposeful signal that grass is forbidden. Predicting observed actions accurately and inferring the intended reward accurately are different evaluation goals.[3]
Here, “literal” describes an interpretation of behavior. In Should Robots be Obedient?, obedience instead means executing an order. The distinction matters: a learner can use a literal behavior model and still choose to depart from an order. The 2017 simulations show that omitted reward features can undermine the apparent benefit of doing so.[4]
What a correction must be able to change
The 2026 one-round result gives deliberate teaching a favorable role under a specially matched human model and available signals.[6] The earlier mismatch results concern what happens when that match fails. They do not contradict a conditional theorem; they identify evidence needed before applying it elsewhere.[3]
The editorial connection to Outer Alignment and Corrigibility is that a correction may challenge the model itself: which preferences exist, what the person knows or whether an action was meant as instruction. Updating weights inside an unchanged model can leave that challenge unanswered. From objectives to accountable control connects this interpretive problem with preserving intervention and informed responsibility.
Changing preferences versus reacting to the interaction
The 2026 online-learning result allows preferences to change, provided their sequence is fixed before interaction begins. Its guarantees also specify the human’s feedback and learning algorithm. Those assumptions differ from a person revising their preference because of the assistant’s earlier action. Appendix B’s adaptive-preference counterexample explains a mathematical boundary; assessing a real interaction still requires checking how its preferences, feedback and behavior are represented.[2]
When a physical correction falls outside the objective space
Learning under Misspecified Objective Spaces gives a concrete robot example: a push intended to increase distance from a person also changes table distance. If the first feature is absent, updating the second can teach a preference the person never intended. The method tests whether a lower-effort correction could have achieved the known feature changes, then reduces the update when those changes seem incidental. “Irrelevant” here describes the input’s fit to the model, not the importance of the human concern.[1]
Its small manipulation study supports conservative learning under the tested conditions. It does not discover the omitted feature. The editorial distinction is between noticing that the current interpretation is inadequate, avoiding damage from that interpretation and building a representation that can express the correction.[1]
Sources
- Learning under Misspecified Objective Spaces · Source record src-218 · Back to claim ↑1 ↑2
- Provably Optimal Learning Algorithms for Assistance Games · Source record src-209 · Back to claim ↑1
- Literal or Pedagogic Human? Analyzing Human Model Misspecification in Objective Learning · Source record src-207 · Back to claim ↑1 ↑2 ↑3
- Should Robots be Obedient? · Source record src-208 · Back to claim ↑1 ↑2
- Cooperative Inverse Reinforcement Learning · Source record src-204 · Back to claim ↑1
- Corrigible Assistance in One Round: Pragmatic-Pedagogic Best Response · Source record src-206 · Back to claim ↑1
Pages that link here
- From objectives to accountable control note
- Learning under Misspecified Objective Spaces paper
- Literal or Pedagogic Human? Analyzing Human Model Misspecification in Objective Learning paper
- Outer Alignment concept
- Provably Optimal Learning Algorithms for Assistance Games paper
- Should Robots be Obedient? paper
Last updated 2026-10-10