When Pooled AUC Is Not Early Warning
A reproducible audit of SAFE-style failure detection on OpenVLA and LIBERO
Strong pooled ROC-AUC does not establish timely warning on an unseen task. This audit examines 1,000 independently collected rollouts across the ten libero_10 tasks, with task-held-out evaluation, conformal calibration, and fixed-horizon checks.
Read the final report Reproduce the analysis
What the evidence shows
| Evaluation | MLP | LSTM |
|---|---|---|
| Mean pooled early-window AUC, original three splits | 0.747 | 0.755 |
| Equal-task early-window AUC, ten-fold LOTO | 0.672 | 0.652 |
| Failure catch at ≤5% realized FPR, earliest-stop LOTO | 8.5% | 6.9% |
At those tested low-FPR operating points, effective mean alarm positions are 95.5% and 95.7% of the evaluation window, counting missed failures as alarm-at-end. The 50/100/148-step sensitivity retains the negative warning conclusion. These findings concern this reproduction; they do not establish that SAFE fails for every policy or benchmark.

Inspect and reproduce
The public score bundle reproduces all eight primary functional-conformal result tables without raw hidden states or trained checkpoints.
- Data and reproduction instructions
- Score-bundle format and integrity contract
- Collection and training provenance
- Source code and downloadable artifacts
The maintainer collected the rollouts using the same OpenVLA weights and seeds as SAFE’s authors. The provenance record distinguishes that confirmation from recovered metadata and the limits of full experiment reproduction.