Experiment provenance and reproduction scope
Updated 2026-09-20. The public score bundle supports replay of the primary functional-conformal results, including 50/100/148-step sensitivity, without raw hidden states or detector checkpoints. Collection and training provenance are described below so that metric replay is not mistaken for bitwise reproduction of the entire experiment.
Rollout collection
The maintainer confirms that these rollouts were collected independently using the same OpenVLA weights and seeds as SAFE’s authors. The raw data are backed up on Google Drive; no public backup URL is part of this release.
SAFE’s adapted OpenVLA collection recipe specifies openvla/openvla-7b-finetuned-libero-10 for libero_10, with center cropping and hidden-state output. Its collector defaults to seed 7, 50 trials per task, one action sample per step, and 10 settling steps; the libero_10 action horizon is 520. These are verified upstream recipe details, consistent with the maintainer’s account, rather than an archived command for each local shard. Collector seed 7 is distinct from probe-training seeds and conformal-split seeds.
The original per-shard launch commands, immutable policy-weight revision, collection-time source revision, and complete simulator environment were not recovered. The local multilayer feature-capture changes also exceed the upstream recipe’s last-layer export. Use the backup for exact source inputs; the upstream recipe alone is not a claim to regenerate identical tensors.
The two retained shards contain 500 rollouts each, 50 per task per shard. The union has 527 successes and 473 failures. Failure lengths are all 520; success lengths range from 148 to 513. Rollout identities include the shard alias because task and episode numbers recur across shards.
Task mapping
Saved metadata confirms libero_10. One rollout per task from each prepared shard was inspected; all 20 metadata samples agree on these descriptions. Their relative paths and hashes are included in the portable provenance record.
| Task ID | Saved task description |
|---|---|
| 0 | put both the alphabet soup and the tomato sauce in the basket |
| 1 | put both the cream cheese box and the butter in the basket |
| 2 | turn on the stove and put the moka pot on it |
| 3 | put the black bowl in the bottom drawer of the cabinet and close it |
| 4 | put the white mug on the left plate and put the yellow and white mug on the right plate |
| 5 | pick up the book and place it in the back compartment of the caddy |
| 6 | put the white mug on the plate and put the chocolate pudding to the right of the plate |
| 7 | put both the alphabet soup and the cream cheese box in the basket |
| 8 | put both moka pots on the stove |
| 9 | put the yellow and white mug in the microwave and close it |
Recovered probe-training settings
Local W&B launch metadata and configurations, checkpoint experiment fields, and training summaries identify the following settings. The portable record retains selected configuration fields, source-file hashes, CLI arguments with tracking options removed, and the selected epoch per fold. Machine-specific paths, account details, and host metadata are omitted.
| Setting | Primary MLP | Primary LSTM |
|---|---|---|
| Input | Layer 32, last action token, 4,096 features, full temporal rate | Same |
| Architecture | Two linear layers, hidden width 256, sigmoid increments and cumulative sum | One LSTM layer, hidden width 256, sigmoid scalar head |
| Training loss | SAFE cumulative loss, threshold disabled | Per-timestep BCE |
| Optimizer / learning rate | Adam / 0.0001 | Adam / 0.0001 |
| Explicit non-bias L2 coefficient | 0.01 | 1.0 |
| Epoch budget | 1,000; patience-based early stopping disabled | Same |
| Saved state selection | Minimum validation loss | Same |
| Batch / accumulation | 64 / 8 batches (up to 512 examples per update; final group can be smaller) | Same |
| Dropout / gradient clipping | 0 / disabled | Same |
| Training device recorded | Apple MPS | Apple MPS |
| LOTO folds / optimizer seeds | Task IDs 0–9 / seeds 0–9, one seed per fold | Same |
| Seen-task split shuffle seed | 0 in every LOTO fold | Same |
| LOTO train / validation / test | 540 / 360 / 100 | Same |
The original pooled comparison uses separate checkpoints with task-split and optimizer seeds 0, 1, and 2. Its train/validation/test sizes are 420/280/300; the held-out task sets are {0, 3, 5}, {5, 7, 8}, and {6, 8, 9}. The other settings above are shared. The primary calibration sweeps seeds 0–9 after training; those are not 10 independent optimizer runs per held-out task.
The saved run metadata names Git revision 8bdb598702e42af0d44bdc5555a8b7c1f8917f4a. This records HEAD rather than a clean-tree attestation: it does not fully identify training-time working-tree changes. The current lock file is the tested release/replay environment, not a recovered collection or historical training environment. Exact output identity is supported by the published score bundle and checkpoint hashes, not promised for fresh optimization on another device.
Training reconstruction commands
First prepare and verify both full-rate layer-32 roots as described in DATA.md. Build a dedicated ordered training union so unrelated data under data/rollouts cannot enter discovery. The prepared artifacts have the homogeneous three-dimensional schema expected by the training loader. Use a fresh union directory containing only these two shards:
mkdir -p data/processed/primary_training/shard0 data/processed/primary_training/shard1
cp -Rl data/processed/primary_l32/openVLA-last-layer/. data/processed/primary_training/shard0/
cp -Rl data/processed/primary_l32/openVLA-last-layer-2/. data/processed/primary_training/shard1/
uv run python scripts/train_openvla_layer_selection.py \
--root data/processed/primary_training \
--output-dir reproduced/training/mlp_loto \
--model-type mlp --layers 32 --loto --loto-seed 0 \
--lr 1e-4 --lambda-reg 1e-2 --optimizer adam --epochs 1000 \
--batch-size 64 --gradient-accumulation-steps 8 \
--token-pool last --hidden-dim 256 --dropout 0 --grad-clip 0 \
--seen-train-ratio 0.6 --device mps
uv run python scripts/train_openvla_layer_selection.py \
--root data/processed/primary_training \
--output-dir reproduced/training/lstm_loto \
--model-type lstm --layers 32 --loto --loto-seed 0 \
--lr 1e-4 --lambda-reg 1 --optimizer adam --epochs 1000 \
--batch-size 64 --gradient-accumulation-steps 8 \
--token-pool last --hidden-dim 256 --dropout 0 --grad-clip 0 \
--seen-train-ratio 0.6 --device mpsThese reconstruct the recorded configuration with explicit, isolated inputs and outputs; they are not the original shell transcript. The layer-selection entrypoint disables patience-based early stopping by default and saves the minimum-validation-loss state. Replace mps with cpu or cuda for another device, acknowledging numerical differences. For the original three-split pooled design, replace --loto --loto-seed 0 with --seeds 0,1,2 --skip-topk --split-mode task --unseen-task-ratio 0.3 and use separate output directories. Compare the embedded splits and source inventory against the published contracts before interpreting any new run as a replication.
The remaining inference and replay commands are in DATA.md. Secondary multilayer training and analysis commands are in the technical report.