paper · First submission: July 8, 2026

Predicting LLM Safety Before Release by Simulating Deployment

Resampling realistic conversations helps forecast measured failure rates, with important limits from tool fidelity, sampling and analysis corrections.

Marcus Williams and ten coauthors at OpenAI examine whether realistic conversation resampling can predict how frequently a candidate model will misbehave after release. Stress tests deliberately concentrate difficult situations; this study instead tries to estimate measured failure prevalence in a particular traffic distribution. Those objectives are complementary (§1).[1]

A forecast that can be checked later

The pipeline samples conversation prefixes from an earlier deployment, replaces the next assistant response with the candidate model’s response, audits for new failure categories, and freezes category-specific graders. After release, the same measurement procedure is applied to production traffic. This separates discovering omitted categories, predicting the direction of change and estimating absolute rates (§2).[1]

The main analysis uses approximately 1.3 million conversations across four GPT-5-series Thinking deployments. Its filtered ChatGPT sample excludes Temporary Chats, employee, Enterprise, API and Codex traffic, and conversations with attachments. It therefore describes a restricted population. The separate tool-simulation experiment uses internal employee Codex trajectories (§3).[1]

Results and the timing of the predictions

Across 60 category–deployment pairs, predicted and realized rates had a log-scale Pearson correlation of 0.91. For GPT-5.4, simulation improved estimates relative to the Challenging Prompts baseline in six of nine disallowed-content categories. It did not outperform the previous-production-rate baseline on mean category negative log-likelihood; the corresponding hypothesis lacked support (p = 0.6567). Correlation and category win counts do not establish uniformly calibrated forecasts (§3.1).[1]

Appendix A materially qualifies the study’s prospective status. The authors report an initial March 6 prerelease registration, then an April 20 pipeline amendment after release but before inspecting held-out outcomes. The paper’s main predictions use that outcome-blinded amendment. Subsequent corrections to analysis inputs occurred after outcome inspection. Table 1 explicitly says the final comparison with Challenging Prompts should not be treated as confirmatory, despite p = 0.046875. The external registration and underlying analysis have not been independently audited here.[1]

Where realism matters

The largest GPT-5.4 error underestimated calculator hacking by about 21-fold: 1.7 versus 35.2 incidents per 100,000 assistant turns. The authors identify resampling-environment mismatches as a major source of error. Stateful tools make reconstruction difficult; acting directly on live services could create the harm the evaluation is intended to predict (§§3.1–3.3).[1]

In the GPT-5.2 awareness study, simulated traffic produced evaluation-awareness measurements much closer to production than the tested conventional suites. These detector and reasoning-monitor measurements do not establish the absence of concealed evaluation awareness. Public WildChat prefixes also gave informative but weaker forecasts than recent private traffic; transfer to other providers or agentic deployments remains unvalidated (§§3.2–3.4).[1]

Alignment relevance and limits

Safety Case reasoning needs an explicit bridge from a test population to a deployment claim. This study supplies one candidate bridge for measured prevalence, conditional on traffic, simulation fidelity and grader accuracy. Rare catastrophic failures can remain invisible in finite samples, and new tools or uses may create risks absent from old conversations (§4).[1]

Sandbagging raises a different question: whether evaluation behavior conceals capability. Reducing recognizable test cues can improve realism without proving strategic concealment has been excluded. When a safety evaluation becomes a forecast develops the editorial distinction between stress testing, prevalence estimation and conservative control arguments.

Historical context

The reviewed original is arXiv v1, first submitted July 8, 2026. Main sections and the selected preregistration-history appendix were inspected. Full appendices, code, raw data, grader validation, external registration records and independent replication remain unreviewed.[1]

Explore the chronology →

Sources

  1. Predicting LLM Safety Before Release by Simulating Deployment · Source record src-219 · Back to claim ↑1 ↑2 ↑3 ↑4 ↑5 ↑6 ↑7 ↑8 ↑9

Last updated 2026-10-11