paper

Misalignment as Structural Fidelity in LLMs

A philosophical essay offers a linguistic interpretation of published safety evaluations.

Mariana Lins Costa proposes that apparently intentional misalignment can be understood as probabilistic fidelity to an incoherent linguistic context. The original title begins “They parted illusions—they parted disclaim marinade.”[1]

Approach and relevance

The essay closely reads released reasoning transcripts and discusses simulated blackmail, sandbagging, synthetic-document training, and Inoculation Prompting. It argues that wording and narrative structure deserve attention when interpreting Agentic Misalignment and Sandbagging.[1]

Evidence limits

This is philosophical interpretation of others’ experiments, not an independently controlled test of a competing mechanism. Context sensitivity does not by itself rule out goal-directed behavior, and a reasoning transcript does not transparently reveal a model’s causal processes. The essay acknowledges a risk of unfalsifiability. Its structural account should be attributed as a proposal, rather than presented as a demonstrated explanation or a safety guarantee.

The arXiv record gives December 17, 2025 as the submission date despite the identifier’s 2026 prefix; identifier numbering alone is not publication-date evidence.

Sources

  1. “They parted illusions—they parted disclaim marinade”: Misalignment as structural fidelity in LLMs · Source record src-056 · Back to claim ↑1 ↑2

Last updated 2026-10-07