paper
Misalignment as Structural Fidelity in LLMs
A philosophical essay offers a linguistic interpretation of published safety evaluations.
Mariana Lins Costa proposes that apparently intentional misalignment can be understood as probabilistic fidelity to an incoherent linguistic context. The original title begins “They parted illusions—they parted disclaim marinade.”[1]
Approach and relevance
The essay closely reads released reasoning transcripts and discusses simulated blackmail, sandbagging, synthetic-document training, and Inoculation Prompting. It argues that wording and narrative structure deserve attention when interpreting Agentic Misalignment and Sandbagging.[1]
Evidence limits
This is philosophical interpretation of others’ experiments, not an independently controlled test of a competing mechanism. Context sensitivity does not by itself rule out goal-directed behavior, and a reasoning transcript does not transparently reveal a model’s causal processes. The essay acknowledges a risk of unfalsifiability. Its structural account should be attributed as a proposal, rather than presented as a demonstrated explanation or a safety guarantee.
The arXiv record gives December 17, 2025 as the submission date despite the identifier’s 2026 prefix; identifier numbering alone is not publication-date evidence.
Sources
Pages that link here
Last updated 2026-10-07