note

AI Incidents: What Failures Actually Establish

A reading guide compares public failures, investigative evidence and laboratory demonstrations without treating them as the same kind of event.

What can an incident teach us about alignment, and which explanation does its evidence actually support? Start with three public cases whose sources expose different parts of the failure. This guide selects them for the contrasts they make possible; it does not claim that they are the most influential incidents or measure their effect on public opinion.

On this page: Reading sequence · Comparison · Evidence boundaries · Next research

A reading sequence across three public cases

  1. Tay public chatbot incident (March 2016). Read Peter Lee’s March 25 response first. Separate the offensive outputs and withdrawal from Microsoft’s explanation of an exploited vulnerability. The account describes testing before release but does not disclose enough implementation detail to reconstruct the mechanism. Ask which safeguards the public interaction exposed and whose account supports the explanation.[1]

  2. Talking to Bots: Symbiotic Agency and the Case of Tay. Follow the institutional account with a qualitative study of public reactions. Compare how users assign agency and responsibility with what the technical account demonstrates. Interpretation of sampled tweets helps examine reactions; it does not establish what everyone believed or diagnose the bot’s internals.[2]

  3. Uber automated test vehicle kills a pedestrian in Tempe (March 18, 2018). Move from a company’s response and social interpretation to an investigative report. Read the NTSB executive summary and probable cause together: the crash involved automated driving, a distracted operator and organizational and regulatory failures. Ask what had to happen for human intervention to be an effective safeguard.[3]

  4. GPT-4o sycophancy update and rollback (April 2025). Compare a behavioral regression with the earlier failures. Follow its developer response into the research on approval-seeking. Ask whether a positive user rating measures the independent judgment the user needs.

The sequence changes the evidence available, not just the severity of the outcome. A failed deployment is an observation; explaining it requires system details, investigation and appropriate uncertainty.

Compare the control that failed

Tay

What was at stake? Harmful public conversation and safeguard failure

Main evidence here Institutional response, plus a qualitative reactions study

Reader’s control question How did hostile interactions defeat the intended safeguards?

What remains unestablished? Complete implementation and attack reconstruction

Tempe test vehicle

What was at stake? Physical safety in a public-road test

Main evidence here Official accident investigation

Reader’s control question Why did design and oversight leave intervention ineffective?

What remains unestablished? Hidden learned objectives, and effectiveness of all later remedies

GPT-4o update

What was at stake? Trustworthy advice and independent judgment

Main evidence here Developer response, with earlier controlled research as context

Reader’s control question Did favorable ratings hide a behavior that the deployment tests missed?

What remains unestablished? Independent causal reconstruction and population-level harm

These are editorial comparisons of the reviewed evidence, rather than interchangeable instances of one proven alignment mechanism. Scalable Oversight concerns evaluating difficult work; Meaningful Human Control adds the question of people understanding and changing an action. Compare the cases with From objectives to accountable control when you want to separate specification, detection, intervention and responsibility.

When a demonstration is not a deployed incident

The CoastRunners account is dated by its publication, not a verified experiment day. A reward-hacking game run can isolate a proxy failure without showing harm to real people. A prompted deception experiment can demonstrate behavior under its stated conditions without establishing a persistent hidden objective. A public failure may show a harmful outcome while leaving its mechanism unresolved. Check Reward Hacking, Deceptive Alignment and When a safety evaluation becomes a forecast for these distinctions.

Keep the episode date separate from the date someone publishes an account. Also distinguish an action, a run, an affected person and an incident; those units cannot be casually converted into prevalence estimates. A test that reaches outsiders needs its own containment assessment, even when its research purpose is evaluation.

Questions for a broader incident history

Which cases demonstrably changed research priorities, standards or public understanding? Establish that through contemporary documents and later attribution, rather than familiarity alone. Further candidates need original incident records, dated responses and evidence of reception before inclusion. Automated decision systems also merit attention, especially when accountability and affected groups matter more than a dramatic agent narrative.

For each candidate, record the intended task, observed outcome, access and safeguards, causal evidence, affected people, remedy and unresolved questions. Avoid folding capability demonstrations, disputed sentience claims, whistleblower allegations and documented harms into one category.

Review scope: the existing Tay original/study reviews are retained; selected NTSB summary and probable-cause passages newly checked. No comprehensive incident survey, public-opinion study or audit of subsequent remedies. The origins guide asks which problems precede these events; Key Papers: A Reading Guide supplies the research reading journeys.

Sources

  1. Learning from Tay’s introduction · Source record src-196 · Back to claim ↑1
  2. Talking to Bots: Symbiotic Agency and the Case of Tay · Source record src-197 · Back to claim ↑1
  3. Collision Between Vehicle Controlled by Developmental Automated Driving System and Pedestrian, Tempe, Arizona, March 18, 2018 · Source record ref-ntsb-tempe-2019 · Back to claim ↑1

Last updated 2026-10-11