note
From objectives to accountable control
Why specifying goals, checking behavior, intervening and preserving human responsibility require different evidence.
An AI system can follow a written objective, receive a favorable evaluation and still leave people unable to govern its effects. This note offers an editorial thread through the encyclopedia: goal specification, evidence for control and informed responsibility answer different questions. Progress on one does not settle the others.
What should the system do?
The 2015 research agenda distinguishes implementing a specification from choosing a requirement that serves the intended purpose. Its vacuum example makes this concrete: collecting dirt is not equivalent to leaving a clean floor.[4] Outer Alignment asks what the objective represents; Goodhart’s Law distinguishes mechanisms by which optimizing a proxy can weaken its relationship to the goal.[5]
The comparison matters because a more capable optimizer can be better at satisfying the measured target without resolving the choice of target. Choosing whose interests count belongs alongside improving the measurement.
Why would correction remain possible?
Corrigibility adds a requirement that a well-chosen initial objective cannot replace: preserving the ability to correct mistakes later. Soares and colleagues show how a shutdown rule can work after a button press while the surrounding incentives still undermine it, including when the agent creates successor software.[9]
Cooperative Inverse Reinforcement Learning makes teaching and learning part of assisting a human. The Off-Switch Game then treats a human’s decision as information about an uncertain objective. Under its assumptions, waiting can improve the shared outcome. Yet uncertainty is not a universal remedy: a system can learn within a flawed rule and still resist correction of that rule. The two papers examine different parts of the problem, rather than establishing that a learning system is corrigible in every respect.[11][10][9]
A further distinction concerns what an intervention should teach. Safely Interruptible Agents treats interruption as an imposed policy and asks how to avoid a lasting bias in learning the task. The off-switch game instead makes the human decision informative about the task’s value. These approaches serve different purposes: an emergency stop may need to override execution, while a revised instruction may need to change what the assistant believes it should do.[12][10]
Interpretation is another condition on that second route. The 2026 one-round assistance result shows how treating a human action only as task execution can leave goals ambiguous. Purposeful signaling overcomes that ceiling in a restricted game with distinct, cost-free signals for every modeled goal. The editorial implication is to examine whether a person can communicate a correction through the available interface, and whether the assistant’s human model makes it legible. The theorem does not supply that evidence for an arbitrary interface or person.[13]
A model can also misunderstand the correction. Human Model Misspecification separates what a person values from how they express it. A learner might mistake an ordinary action for deliberate teaching, or omit a preference that makes an order sensible. Better prediction of behavior does not necessarily mean better inference of the goal.[2][3] In the records example, the editorial question is whether preservation of irreplaceable information can enter the assistant’s model at all, alongside whether it accepts the revised instruction.
For the records example below, this adds an editorial question: can the assistant pause and accept a revised requirement even when its current estimate favors proceeding? That differs from merely recognizing that the proposed deletion carries risk.
Which incentives reach the actual system?
The Basic AI Drives asks why preserving operation or obtaining resources can help many objectives. Optimal Policies Tend to Seek Power sharpens part of that idea into a conditional optimal-policy result, while its reviewed revision warns that learned policies can differ substantially.[18][19] Instrumental Convergence therefore adds two editorial questions: what makes the intermediate action useful, and what evidence shows the deployed system responds to that incentive?
In the records example, retaining access could help finish the reorganization. That possibility alone does not show the assistant will resist revoking its access. Conversely, a claim that the assistant lacks a human survival instinct does not establish that intervention will remain available. Connecting a theory of incentives to observed behavior is a separate evidential step.
What would detect a failure and change the action?
A control claim needs a response as well as a detector. Attack-selection research shows why the opportunities an attacker searches and the information it receives belong in an evaluation’s description.[6] A good result against one attack-selection procedure need not bound another.
Safety Case connects a claim to evidence and assumptions. AISI’s cyber-evaluation incident also makes the evaluation’s own containment relevant: the institute reports changes to network restrictions, monitoring and security review.[7] Reports that a control exists support a different claim from evidence that it remains effective in a new setting.
Who understands and owns the decision?
Meaningful Human Control adds informed human responsibility to responsiveness to relevant reasons. Santoni de Sio and van den Hoven argue that a person’s nominal role does not establish the needed understanding of capabilities, consequences and responsibility. They also distinguish control from morally acceptable objectives.[8]
This creates a bridge to Machine Ethics and Society-in-the-Loop: which reasons should govern the system, whose interests can challenge the decision, and what makes the assigned human role feasible? These are editorial connections across different frameworks, rather than a claim that their authors share one theory.
A question linking the layers
Consider an editorial example: an assistant receives approval to reorganize records, but someone discovers that a proposed step will erase irreplaceable information. The goal needs a constraint, the evaluation needs to notice the consequence, an intervention needs to take effect, and the responsible people need to understand the system’s limits. These are distinct requirements even when they concern one action.
A useful question across the corpus is therefore: what can the people affected understand, challenge and change before an irreversible consequence occurs? Answering it requires examining the agent, the interface and the institution together. It remains a question for investigation, not a completed safety guarantee.
When should an institution allow the activity?
The 2023 pause letter on giant AI experiments asks for a temporary training halt and independent review; the 2023 statement on AI risk identifies a global mitigation priority without prescribing that halt.[14][15] The editorial distinction is between identifying a serious risk and choosing the conditions under which an activity may proceed.
A system’s willingness to accept correction does not settle who may authorize its training or deployment. Likewise, a public appeal does not supply the evidence for a safety claim. Extending the records example to an institution adds questions about who sets an acceptable exposure, who checks the evidence, and who can suspend the activity when the case is inadequate.
Which pathway does the evidence address?
Is Power-Seeking AI an Existential Risk? separates a chain of agent and deployment conditions; Current and Near-Term AI as a Potential Existential Risk Factor instead maps how institutions and information systems could amplify wider hazards.[16][17] The editorial implication is to name the pathway before claiming that a control addresses it.
In the records example, accepting a stop instruction addresses the assistant’s conduct. It leaves a different question: can the institution discover whose records were affected, hear their objections and repair the decision? A safety case should make those responsibilities visible. A public appeal can motivate investigation; choosing an intervention requires evidence about the particular mechanism it is intended to change.
What if the correction has no place in the model?
Learning under Misspecified Objective Spaces separates avoiding an unintended update from understanding the missing preference. A robot can reduce learning from a push it cannot explain without learning the concern that motivated the push. That measured benefit addresses one failure in correction, while leaving model expansion unresolved.[1]
For the records example, preserving the current plan after an unexplained correction would still be insufficient if the plan erases irreplaceable information. The editorial question is what happens after recognition of a model limit: can execution pause, can a person revise the representation, and can the new requirement actually govern the action? The robot study does not test those broader controls; it helps make their distinct roles visible.
Sources
- Learning under Misspecified Objective Spaces · Source record src-218 · Back to claim ↑1
- Literal or Pedagogic Human? Analyzing Human Model Misspecification in Objective Learning · Source record src-207 · Back to claim ↑1
- Should Robots be Obedient? · Source record src-208 · Back to claim ↑1
- Research Priorities for Robust and Beneficial Artificial Intelligence · Source record src-191 · Back to claim ↑1
- Categorizing Variants of Goodhart's Law · Source record ref-e178cbf073c0 · Back to claim ↑1
- Attack Selection Reduces Safety in Concentrated AI Control Settings against Trusted Monitoring · Source record src-189 · Back to claim ↑1
- Building a more secure environment for evaluating dangerous capabilities · Source record src-184 · Back to claim ↑1
- Meaningful Human Control over Autonomous Systems: A Philosophical Account · Source record src-201 · Back to claim ↑1
- Corrigibility · Source record src-202 · Back to claim ↑1 ↑2
- The Off-Switch Game · Source record src-203 · Back to claim ↑1 ↑2
- Cooperative Inverse Reinforcement Learning · Source record src-204 · Back to claim ↑1
- Safely Interruptible Agents · Source record src-205 · Back to claim ↑1
- Corrigible Assistance in One Round: Pragmatic-Pedagogic Best Response · Source record src-206 · Back to claim ↑1
- Pause Giant AI Experiments: An Open Letter · Source record src-210 · Back to claim ↑1
- Statement on AI Risk · Source record src-211 · Back to claim ↑1
- Is Power-Seeking AI an Existential Risk? · Source record src-214 · Back to claim ↑1
- Current and Near-Term AI as a Potential Existential Risk Factor · Source record src-215 · Back to claim ↑1
- The Basic AI Drives · Source record src-216 · Back to claim ↑1
- Optimal Policies Tend to Seek Power · Source record src-217 · Back to claim ↑1
Pages that link here
- Corrigibility concept
- Human Model Misspecification concept
- Instrumental Convergence concept
- Learning under Misspecified Objective Spaces paper
Last updated 2026-10-10