paper · First submission: February 19, 2025
Safe Learning Under Irreversible Dynamics via Asking for Help
A 2026 theoretical result combines mentor queries and local generalization to approach mentor performance without resetting after irreversible errors, under strong assumptions.
- Received: September 2025
- Revision: April 2026
- Online publication: June 2026
- Revision: September 11, 2026
Benjamin Plaut, Juan Liévano-Karim, Hanlin Zhu and Stuart Russell study learning when a bad action can cause lasting damage. Their JMLR paper develops a conditional theoretical guarantee for an agent that asks a mentor for an action when needed and transfers guidance between sufficiently similar states.[1]
Counting damage caused while learning
The setting permits infinite-state, non-communicating Markov decision processes: the learner cannot necessarily return to earlier states and cannot reset. Regret compares the expected total rewards of learner and mentor on their own trajectories from the same starting state. Evaluating both policies only on the learner’s states could ignore damage from entering a permanently bad state; Figure 2 illustrates this with a choice between two absorbing outcomes.[1]
The mentor can be suboptimal. Approaching mentor performance therefore does not establish optimal behavior, correct human values or absolute safety.[1] Reinforcement Learning: A Survey provides earlier context on reward criteria and penalties incurred during learning; Concrete Problems in AI Safety discusses safe exploration as an accident-risk problem.
Asking for help and transferring guidance
The argument uses three reductions: full-feedback online learning to active learning, active learning to avoiding catastrophe with a mentor, and catastrophe avoidance to low regret in multichain MDPs. The illustrative algorithm combines occasional learning queries with additional queries when a proposed action lacks nearby mentor-demonstrated support (Algorithm 3). Guidance is available before the action is executed.[1]
Local generalization constrains how much transition outcomes can change when reusing the mentor’s action from a nearby state. The result also relies on a known, realizable mentor-policy class with finite learning complexity: finite VC dimension with smooth state distributions, or finite Littlestone dimension in the fully adversarial case. The concrete construction uses binary actions. Theorem 10 gives sublinear expected reward regret and a mentor-query bound with a sublinear time factor; its query bound also depends on the expected diameter of the encountered states. This is not an unconditional guarantee for arbitrary expanding environments.[1]
For Scalable Oversight, the editorial connection is to reducing the frequency of external help while accounting for irreversible consequences. The theorem assumes a mentor and a suitable similarity channel; it does not demonstrate that a human can supervise every more-capable system or detect unfamiliar states reliably.
Practical limits
Section 7 says the algorithm is not ready for practical use. Computing its policy cover is expensive; bounds become barely sublinear in high dimensions. It requires exact distance computation, likened to a perfect out-of-distribution detector, full state observability and knowledge of the policy class. Learning from an initial offline dataset, imperfect detection and unavailable mentors remain future directions.[1]
Historical context and review limits
The first arXiv submission was February 19, 2025, which supplies the chronology date. The JMLR original records September 2025 submission, April 2026 revision and June 2026 publication. The September 11, 2026 arXiv revision is a separate version; no full version comparison was performed. Selected original sections 1–3, 4.3, 5.2 and 6–7, Algorithm 3 and rendered pages 4 and 17 were checked. Full reduction proofs, appendices and cited studies were not independently audited; this review does not establish practical deployment performance.[1]
Sources
- Safe Learning Under Irreversible Dynamics via Asking for Help · Source record src-178 · Back to claim ↑1 ↑2 ↑3 ↑4 ↑5 ↑6 ↑7
Pages that link here
- Concrete Problems in AI Safety paper
- Distributional Shift concept
- Outer Alignment concept
- Safety Case concept
- Scalable Oversight concept
- Stuart Russell person
Last updated 2026-10-09