paper

AI Deception: Risks, Dynamics, and Controls

A survey maps deceptive behavior, evaluation methods, and possible mitigations.

The project team, including lead contributors Boyuan Chen and Jiaming Ji, organizes research around a deception cycle: emergence through incentives, capabilities, and context; treatment through detection, evaluation, and mitigation.[1]

Contribution and relevance

Its functional definition emphasizes signals that induce false beliefs, change a receiver’s behavior, and advantage the sender. This avoids treating consciousness or subjective intention as prerequisites. The taxonomy connects Sandbagging, Alignment Faking, and Reward Hacking while distinguishing misleading outputs from ordinary mistakes.[1]

Evidence limits

This is a literature synthesis and conceptual framework, not a new experiment establishing a universal mechanism. Its executive-summary assertions that deception cannot be removed without damaging intelligence, or grows exponentially with capability, are stronger than the surveyed evidence establishes. Use the survey to locate original studies; do not cite those assertions as settled facts. A measured deceptive behavior does not establish the inevitability of deception in every capable system.

Sources

  1. AI Deception: Risks, Dynamics, and Controls · Source record src-002 · Back to claim ↑1 ↑2

Pages that link here

Last updated 2026-10-07