When Pooled AUC Is Not Early Warning

A reproducible audit of SAFE-style failure detection on OpenVLA and LIBERO

Strong pooled ROC-AUC does not establish timely warning on an unseen task. This audit examines 1,000 independently collected rollouts across the ten libero_10 tasks, with task-held-out evaluation, conformal calibration, and fixed-horizon checks.

Read the final report Reproduce the analysis

What the evidence shows

Evaluation MLP LSTM
Mean pooled early-window AUC, original three splits 0.747 0.755
Equal-task early-window AUC, ten-fold LOTO 0.672 0.652
Failure catch at ≤5% realized FPR, earliest-stop LOTO 8.5% 6.9%

At those tested low-FPR operating points, effective mean alarm positions are 95.5% and 95.7% of the evaluation window, counting missed failures as alarm-at-end. The 50/100/148-step sensitivity retains the negative warning conclusion. These findings concern this reproduction; they do not establish that SAFE fails for every policy or benchmark.

Failure catch rate versus realized false-positive rate, with color showing effective alarm position.

Inspect and reproduce

The public score bundle reproduces all eight primary functional-conformal result tables without raw hidden states or trained checkpoints.

The maintainer collected the rollouts using the same OpenVLA weights and seeds as SAFE’s authors. The provenance record distinguishes that confirmation from recovered metadata and the limits of full experiment reproduction.