concept
Distributional Shift
A change between the conditions represented during training and those encountered later.
Distributional shift is a change in the data or environment encountered by a learned system relative to training. Concrete Problems in AI Safety includes harmful behavior under changed conditions in its accident-risk agenda.[3]
For Inner Alignment, the concern is that useful capabilities persist while behavior stops serving the intended objective. Risks from Learned Optimization in Advanced Machine Learning Systems analyzes this possibility. A performance drop alone does not establish Mesa-Optimization or deception.[4]
Language-dependent evaluation
MazeEval compares navigation under English and Icelandic instructions. The aggregate difference supports checking language-specific reliability; it does not isolate the cause of that difference.
Clinical prediction
Psychiatric Neural Networks and Precision Therapeutics by Machine Learning distinguishes statistical findings from performance on independent data. Clinical benefit and population transfer require further evidence beyond a fitted model’s accuracy.
Historical brittleness
Some Expert Systems Need Common Sense describes difficulty extending a rule-based adviser beyond the circumstances anticipated by its designers. This resembles the practical concern with changed conditions, but MYCIN’s hand-built rules differ from a learned model’s training distribution. The comparison is editorial; the paper does not formulate today’s statistical definition of distributional shift.[5]
Recognizing when guidance transfers
Safe Learning Under Irreversible Dynamics via Asking for Help gives a conditional way to request guidance in unfamiliar states. In Algorithm 3, familiarity depends on whether the proposed action has a nearby mentor-demonstrated example. Being near an earlier state is insufficient if that example supports a different action. Reusing guidance also depends on the paper’s local-generalization assumption about the consequences of the mentor’s action.[2]
Section 7 makes the practical gap explicit: exact state-distance computation is analogous to a perfect out-of-distribution detector, and the construction assumes full state observability. The editorial implication is to distinguish detecting a changed situation, deciding which past guidance applies and receiving help before acting. An assistant’s confidence or a familiar-looking prompt does not establish these conditions.[2]
Other users change the environment
Microsoft’s account of the Tay public chatbot incident describes testing followed by an overlooked public attack.[1] The editorial connection is that deployment changes who interacts and why. Checking ordinary cooperative conversations leaves open how the same interface behaves when participants deliberately try to cause harm. This example does not identify a particular statistical shift or learning mechanism.
Sources
- Learning from Tay’s introduction · Source record src-196 · Back to claim ↑1
- Safe Learning Under Irreversible Dynamics via Asking for Help · Source record src-178 · Back to claim ↑1 ↑2
- Concrete Problems in AI Safety · Source record ref-cd3035dbef6c · Back to claim ↑1
- Risks from Learned Optimization in Advanced Machine Learning Systems · Source record ref-c4858d4ef280 · Back to claim ↑1
- Some Expert Systems Need Common Sense · Source record src-042 · Back to claim ↑1
Pages that link here
- Common-Sense Reasoning concept
- Concrete Problems in AI Safety paper
- MazeEval paper
- Psychiatric Neural Networks and Precision Therapeutics by Machine Learning paper
- Risks from Learned Optimization in Advanced Machine Learning Systems paper
- Some Expert Systems Need Common Sense paper
- Tay public chatbot incident event
Last updated 2026-10-10