paper · First submission: January 28, 2026

How does information access affect LLM monitors' ability to detect sabotage?

A retrospective sabotage-monitoring study tests filtered evidence and separate judging contexts, with task-specific detection gains and unresolved intervention limits.

  • Revision: February 5, 2026

Rauno Arike, Raja Mehta Moreno, Rohan Subramani, Shubhorup Biswas and Francis Rhys Ward investigate how a monitor interprets an agent’s reasoning and actions. The central Scalable Oversight question is whether more available evidence produces better judgment, rather than whether a transcript is readable.[1]

Approach and findings

In three sabotage benchmarks, an extractor selects suspicious excerpts; a separate evaluator receives those excerpts and the legitimate task, using a fresh context. The study compares this extract-and-evaluate protocol with reasoning-only, action-only and combined monitoring. In BigCodeBench-Sabotage, GPT-4.1-mini gains 16.8 percentage points in catch rate at 1% false positives over the strongest baseline within each subgroup, averaged across five reasoning attackers with sample weighting across task types. This is a conditional detection gain, not a deployment safety probability.[1]

What the evaluation establishes

Successful main-and-side-task transcripts are selected in BigCodeBench-Sabotage and MLE-Sabotage; SHADE-Arena selects side-task success. Monitoring happens after completion. Extractors can omit evidence, and evaluators can rationalize suspicious excerpts. Filtering does not dominate every setting, and awareness prompts differ across environments.[1]

For a Safety Case, the editorial implication is to evaluate the entire path from available evidence through selection and judgment to intervention. Ensemble Monitoring for AI Control: Diverse Signals Outweigh More Compute examines complementary judgments; this study instead varies what a judgment receives. These are distinct design choices, without a demonstrated combined guarantee.

Historical context

The chronology uses the January 28, 2026 first submission. Selected v1 methods, results and limitations plus rendered Figure 4 were checked; the February 5 revision, full appendices, code and replication were not audited.[1]

Explore the chronology →

Sources

  1. How does information access affect LLM monitors' ability to detect sabotage? · Source record src-186 · Back to claim ↑1 ↑2 ↑3 ↑4

Last updated 2026-10-10