paper · First submission: July 29, 2026
Corrigible Assistance in One Round: Pragmatic-Pedagogic Best Response
A 2026 assistance-game result makes goal inference tractable when distinct human signals can communicate every goal without sacrificing task value.
Elle Lazarski and Jaime Fernández Fisac study a class of human–robot assistance games in which a person’s action can resolve the robot’s uncertainty in one time step. This entry reviews the July 29, 2026 arXiv v1 manuscript. It combines a formal result with a simple collaborative block-building example, rather than a deployed robot or human-participant study.[1]
The inference ceiling
A robot may assume that a person acts only to complete their own task. But an action can be optimal under several possible goals, so seeing even perfectly task-optimal behavior need not reveal which goal the person intends. The paper derives a ceiling on the robot’s posterior confidence under this model (§5.1). In its two-goal example, the ceiling is 2b/(1+b), where b is the robot’s prior probability of one goal; it stays below certainty for an interior prior.[1]
The proposed alternative models the person as considering how an action will influence the robot’s response. The robot then interprets the action as a purposeful cue. An action that looks equivalent for task execution alone can become informative through this interaction. This extends the teaching-and-learning structure of Cooperative Inverse Reinforcement Learning.[1][2]
A conditional one-round result
The authors define an action-separable game at a particular state and prior. Each possible goal must have a distinct set of signaling actions, and the signals must preserve the value achievable by a fully informed team: communication has zero task cost. Under these conditions and the limiting rational human and robot models, one round of best responses reaches an optimal equilibrium for the full horizon (Definitions 3–4 and Theorem 1, pp.12–14).[1]
Pragmatic-Pedagogic Best Response (PPBR) checks these conditions and computes the signaling map. Its stated work scales with the product of the goal and action-set sizes given the oracle task values. Those values come from solving the fully informed task; the method does not eliminate that planning cost (§6, Algorithm 1).[1]
Figures 2–3 show how the block-building model’s posterior varies with prior beliefs and rationality parameters. Finite rationality can materially affect inference; the exact equilibrium claim concerns the stated limit. The curves illustrate the model and do not measure real users’ ability to understand or produce its signals.[1]
Historical context
The paper builds on CIRL and pragmatic-pedagogic assistance research. Its contribution narrows the computational problem to a class whose communication structure permits an exact short-horizon solution. It also engages earlier concerns about misspecifying whether humans intend to teach; its proposed responsiveness is a theoretical perspective within the modeled setting (§2).[1]
What corrigibility means here
The proposed robot gives the human a way to convey and achieve each modeled goal despite an initially unfavorable belief. This is a useful form of responsiveness within the game. The broader 2015 correction problem also includes correcting the objective-learning rule and maintaining intervention mechanisms; those properties are not established by this theorem.[1][3]
The result depends on the modeled goals, signaling opportunities, task dynamics and common payoff. Section 7 leaves shared signals that identify more than one goal and a multi-step extension outside scope. The paper’s favorable interpretation of pragmatic inference therefore needs to be read alongside its special game class, rather than as proof that richer human models always improve real-world alignment.[1]
Sources
- Corrigible Assistance in One Round: Pragmatic-Pedagogic Best Response · Source record src-206 · Back to claim ↑1 ↑2 ↑3 ↑4 ↑5 ↑6 ↑7 ↑8 ↑9
- Cooperative Inverse Reinforcement Learning · Source record src-204 · Back to claim ↑1
- Corrigibility · Source record src-202 · Back to claim ↑1
Pages that link here
Last updated 2026-10-10