Primary 1,000-rollout audit
Full inference expects two ordered roots:
data/rollouts/openVLA-last-layer/
data/rollouts/openVLA-last-layer-2/
Each root contains 500 pickle files named task<TASK>--ep<EPISODE>--succ<0|1>.pkl. Together they must contain exactly 100 rollouts for each task 0 through 9. Root order is part of the checkpoint split contract and must not be changed.
Each artifact is a dictionary with:
hidden_states:(steps, action_tokens, layers * hidden_dim)for the first historical schema, or(steps, layers * hidden_dim)for an artifact already pooled to the last action token;hidden_state_layers: saved model-layer identifiers including layer 32;hidden_state_dim_per_layer: 4096;hidden_state_token_index: -1for the pre-pooled schema;- task, episode, and success metadata, either in the dictionary or filename.
Low-space staged preparation
The finalized probes need only the full-rate layer-32 last-token trajectory. When both full source shards cannot coexist on disk, each 500-rollout shard can be restored, normalized, verified, and removed from staging before restoring the next one. The preparation script accepts both historical source schemas and does not change tensor values or temporal resolution:
uv run python scripts/prepare_openvla_layer_subset.py \
--input-root data/staging/openVLA-last-layer \
--output-root data/processed/primary_l32/openVLA-last-layer \
--layers 32 --token-index -1 --temporal-stride 1 \
--expected-rollouts 500 --expected-tasks 0,1,2,3,4,5,6,7,8,9 \
--expected-rollouts-per-task 50 --minimum-length 148
uv run python scripts/prepare_openvla_layer_subset.py \
--input-root data/staging/openVLA-last-layer-2 \
--output-root data/processed/primary_l32/openVLA-last-layer-2 \
--layers 32 --token-index -1 --temporal-stride 1 \
--expected-rollouts 500 --expected-tasks 0,1,2,3,4,5,6,7,8,9 \
--expected-rollouts-per-task 50 --minimum-length 148
uv run python scripts/verify_primary_prepared_data.py \
--root data/processed/primary_l32/openVLA-last-layer \
data/processed/primary_l32/openVLA-last-layer-2 \
--output runs/results_audit/primary_prepared_data_audit.jsonOnly remove a staging shard after its output manifest reports 500 rollouts, 50 per task, a homogeneous [32] schema, temporal stride 1, and minimum length at least 148. Preserve both manifests for release provenance. For inference, substitute the two data/processed/primary_l32/... directories for the raw --root paths below, in the same order. The stride-4 500-rollout multilayer set is not a valid substitute.
The finalized checkpoints are expected under:
runs/openvla_mlp_loto_l32/
runs/openvla_lstm_loto_l32/
Each directory must contain folds 0 through 9, with the exact split indices embedded in each checkpoint. The score-bundle manifest records SHA-256 hashes for all 20 checkpoints.
Generate the public replay bundle and all conformal tables with:
uv run python scripts/verify_primary_checkpoints.py
uv run python scripts/evaluate_functional_cp_loto.py \
--root data/rollouts/openVLA-last-layer data/rollouts/openVLA-last-layer-2 \
--mlp-checkpoint-dir runs/openvla_mlp_loto_l32 \
--lstm-checkpoint-dir runs/openvla_lstm_loto_l32 \
--score-bundle-out docs/results_audit/score_bundle \
--output-dir runs/results_audit/functional_cp_loto_closeout \
--alphas 0.01,0.025,0.05,0.075,0.1,0.15,0.2,0.3 \
--cp-seeds 0,1,2,3,4,5,6,7,8,9 \
--horizon 520 --fixed-horizons 50,100,148 \
--task-bootstrap 10000 --bootstrap-seed 20260716After reviewing the generated tables, publish the eight replay-checked result tables named by scripts/compare_functional_cp_outputs.py into docs/results_audit/. Then replay from the bundle into a separate directory and compare:
uv run python scripts/evaluate_functional_cp_loto.py \
--score-bundle-in docs/results_audit/score_bundle \
--output-dir reproduced/functional_cp_loto \
--alphas 0.01,0.025,0.05,0.075,0.1,0.15,0.2,0.3 \
--cp-seeds 0,1,2,3,4,5,6,7,8,9 \
--horizon 520 --fixed-horizons 50,100,148 \
--task-bootstrap 10000 --bootstrap-seed 20260716 --device cpu
uv run python scripts/compare_functional_cp_outputs.py \
--expected-root docs/results_audit \
--actual-root reproduced/functional_cp_lotoThe fixed horizons were declared before this closeout rerun. The longest, 148, is the observed global minimum rollout length. Every rollout is therefore observed through each horizon without padding or exclusion; timesteps after the selected common observation window are intentionally ignored.