paper · First submission: October 11, 2018

Learning under Misspecified Objective Spaces

A robot can reduce unintended learning by testing whether a physical correction makes sense within its known objective features.

  • Proceedings: October 2018
  • Revision: October 26, 2018

Andreea Bobu, Andrea Bajcsy, Jaime F. Fisac and Anca D. Dragan study learning robot objectives from physical human corrections. Their central distinction is between uncertainty about the weights of known features and uncertainty about whether those features can express what the person wants. A belief over candidate objectives can remain incomplete even when it represents uncertainty accurately (§1).[1]

A correction can teach the wrong thing

Figure 1 gives the problem a concrete form. A person pushes a robot away from their body, but the robot does not represent distance from people. The movement also changes distance from the table, so an ordinary update can mistakenly infer a preference about table distance. Calling the correction irrelevant means that it falls outside the robot’s assumed objective space; it does not mean the person’s concern is unimportant (§1, Figure 1).[1]

Learning conservatively

The method asks whether a lower-effort correction could have produced the same changes in the features the robot knows. If so, those changes may be incidental to a different intention. It estimates a rationality parameter from this efficiency comparison, uses empirically fitted relevance likelihoods, and moderates the objective-weight update accordingly. At the limiting case of inferred irrelevance, it retains its prior weights rather than inventing an explanation inside the known space (§2, Equations 8–14).[1]

This remains a model of human behavior: people are assumed to trade off trajectory cost and effort. A correction that looks inefficient under that model is evidence for caution, not proof of a particular omitted preference.

Evidence from robot interaction

The authors collected offline corrections from 12 people using a seven-degree-of-freedom JACO arm. They varied the robot’s available features when interpreting corrections about table distance, human distance and cup orientation. The estimated rationality parameter was higher for corrections to represented features (§3).[1]

A separate online study recruited 12 campus participants, aged 18–30, ten with technical backgrounds. In a randomized within-subject comparison, participants corrected four cup-carrying tasks under fixed and adaptive learning. For corrections outside the hypothesis space, the adaptive method had lower feature-space regret in the authors’ post-hoc analysis (p = 0.001); inside the space, they found no significant difference (p = 0.9991). Participants also rated the adaptive method more favorably for avoiding unintended learning (§§4.1–4.2, Figure 5 and Table 1).[1]

These are controlled manipulation tasks, not evidence about unrestricted assistants. No significant difference in this small study does not establish general equivalence. The method assumes observed human and object locations and uses hand-designed features and fitted relevance distributions. It reduces mistaken updates without identifying or adding the missing feature; model expansion is explicitly left for further work (§5).[1]

Alignment relevance

Human Model Misspecification includes both an incomplete preference space and an unsuitable interpretation of behavior. Should Robots be Obedient? supplies an earlier simulated example of missing reward features; this paper studies physical correction and a way to limit the resulting damage. The editorial connection to Outer Alignment is that learning an objective still depends on what the representation admits. Its connection to Corrigibility is narrower: moderating a mistaken update is one component of accepting correction, not a guarantee that intervention remains available. From objectives to accountable control develops that distinction.[1]

Historical context

The first arXiv submission was October 11, 2018; arXiv v4 was revised October 26. The paper appeared in CoRL 2018 proceedings, PMLR volume 87, pp.796–805, for the October 29–31 conference. This entry reviews the publisher’s ten-page PDF; full equivalence with the arXiv versions has not been checked.[1]

Explore the chronology →

Sources

  1. Learning under Misspecified Objective Spaces · Source record src-218 · Back to claim ↑1 ↑2 ↑3 ↑4 ↑5 ↑6 ↑7 ↑8

Last updated 2026-10-10