Closeout requirements

This document turns the five publication tasks into verifiable acceptance criteria. The project is complete only when every item below is satisfied in a clean checkout.

Completion semantics

An item is passed only when its listed artifacts exist and its verification command exits zero in a fresh clone. A missing prerequisite is blocked, not waived; a development-only flag such as --allow-missing-score-bundle cannot be used as release evidence. Requirement 1 is deliberately last: the publication commit is made only after requirements 2–5 pass.

Hosting and release status

The GitHub Pages site is built from the report and supporting publication files. Every push to main runs the publication workflow; deployment requires the tests, strict artifact checks, eight-table replay, and generated-site link checks to pass. The workflow records the current release status and deployed URL. build-info.json on the site identifies the deployed source commit.

Build the site locally with uv run python scripts/build_site.py. Only publication files are staged into reproduced/pages/_site; raw rollouts, checkpoints, and local tracking metadata are not site inputs.

Local acceptance record

Local acceptance snapshot (2026-09-20, before publication):

Requirement Status Evidence or blocker
1. Publication commit Pending The publication cleanup is prepared locally; committing and publishing it remain separate release steps.
2. Fixed-horizon sensitivity Passed locally All 50/100/148-step tables replay successfully and the primary report discusses them. The LSTM has no tested point meeting the 5% macro FPR cap at 50 or 100 steps.
3. Compact score release Passed locally Schema-v2 bundle: 9,200 records, 1,000 identities, all 20 checkpoint hashes. All eight result tables match replay.
4. Environment/license/citation Passed locally Fresh frozen environment installed; citation schema and publication checks passed. License and collection/training provenance are included.
5. Tests and CI Passed locally; remote CI pending 33 tests passed, sources compiled, strict verification passed, and Quarto 1.9.37 reproduced identical Markdown. CI enforces these checks and score replay.

Verification on 2026-09-20 used a fresh local Git clone with the intended release edits copied in, its own frozen uv environment (macOS, CPython 3.13.3), and no ignored rollout or checkpoint inputs. Score replay matched all eight tables at the comparison script’s default tolerances. This validates the release files before committing; it does not attest a GitHub Actions run or repeat raw collection, probe training, or hidden-state inference. The configured GitHub runner uses Ubuntu and Python 3.11.

The earlier missing-input blocker has been resolved for the published primary score analysis. DATA.md retains the staged full-rate layer-32 preparation route for maintainers reconstructing inference inputs. Raw collection follows the maintainer-confirmed SAFE weights and seeds; exact launch commands and immutable collection dependencies remain outside the public release. See EXPERIMENT_PROVENANCE.md for the evidence boundary.

1. Publication commit

  • The Quarto source, freshly rendered GitHub Markdown, final analysis code, and every report-linked compact artifact are tracked.
  • Quarto 1.9.37 renders docs/safe_openvla_audit.qmd to the tracked Markdown; CI fails if rendering changes it. Edit the source, not the generated report.
  • Generated raw data, checkpoints, and development runs remain ignored. The two already-tracked historical records under runs/openvla_layers/ remain archived separately from the final audit’s evidence.
  • All local Markdown links resolve and all published CSV/JSON files parse.
  • docs/results_audit/publication_manifest.json is regenerated after every other publication edit and verifies every listed hash.
  • The manifest contains only publication artifacts, with no Finder files or hidden render caches. It must pass in a clean checkout without ignored experiment data.
  • uv run python scripts/verify_publication.py exits zero without a waiver.
  • The final commit leaves no intended publication changes uncommitted; git status --porcelain is empty after the commit.

2. Fixed-horizon sensitivity

  • The ordered primary roots contain exactly 1,000 readable rollouts: 500 per root, 100 per task, with dataset indices 0–999 matching the checkpoint split contract. Each finalized checkpoint family contains folds 0–9; every fold has 540 train, 360 validation, and 100 held-out-test indices, and MLP/LSTM split mappings are identical.
  • Evaluate the finalized layer-32 MLP and LSTM checkpoints under all ten LOTO folds at common horizons 50, 100, and 148 steps. These values were selected before the rerun; every validation and test trajectory must have at least 148 observed scores. The evaluator must fail instead of padding, extending, or excluding a trajectory at a fixed horizon.
  • At each horizon, rerun SAFE functional conformal calibration for the existing alpha grid 0.01,0.025,0.05,0.075,0.1,0.15,0.2,0.3 and CP split seeds 0–9.
  • Aggregate CP seeds within held-out task, then report equal-task macro means and task-bootstrap intervals.
  • Publish fixed-horizon raw, per-task, macro, and operating-point tables. Their expected row counts are 4,800, 480, 48, and 36 respectively.
  • Update the technical report with the fixed-horizon low-FPR results and state explicitly whether the primary conclusion changes.

The authoritative generation command is the raw-inference command in DATA.md. Passing this item requires that its output manifest records all three fixed horizons and all 20 checkpoint hashes.

The checkpoint-only half of the prerequisite is independently checked with:

uv run python scripts/verify_primary_checkpoints.py

3. Compact score release and provenance

  • Publish validation and held-out-test score trajectories for every MLP/LSTM LOTO fold in a non-pickle bundle. The bundle must contain exactly 9,200 model/fold/split records and 1,000 unique rollout identities. Every model/fold has 360 validation and 100 test records; held-out test indices cover 0–999 exactly once per model, and fold k tests only task k.
  • Include rollout metadata, split role, task/outcome labels, lengths, and task_min_step, but no machine-local absolute paths or raw hidden states.
  • Include SHA-256 checksums for bundle files and source checkpoints, a source inventory digest, schema version, generation command, and generator hashes.
  • The bundle checkpoint hashes and normalized split digest must match the tracked primary_checkpoint_audit.json; generator hashes must match the current inference, model, dataloader, and conformal source files.
  • A verifier must reject modified files, invalid offsets, unsafe paths, and incomplete metadata, overlapping validation/test indices, inconsistent cross-fold rollout identities, or divergent MLP/LSTM split mappings.
  • Replaying from the bundle must reproduce both the original functional-CP tables and the fixed-horizon tables without raw latents or checkpoints. CI performs the replay and compares eight published result tables with scripts/compare_functional_cp_outputs.py.

4. Environment, license, and citation

  • Direct runtime dependencies live in pyproject.toml; exact transitive versions live in uv.lock.
  • uv sync --frozen --extra dev creates the tested environment.
  • The repository contains an explicit software license and a valid citation file. Third-party papers/code/data remain governed by their own licenses.

Acceptance commands:

uv sync --frozen --extra dev
uv run --frozen --extra dev cffconvert --validate --infile CITATION.cff

5. Tests and continuous integration

  • Synthetic tests cover rollout schema loading, masking, LOTO split isolation, functional-CP calibration, fixed-horizon no-extension behavior, score-bundle round trips/checksums, pooled-AUC decomposition, and multilayer controls.
  • CI installs from the frozen lock, compiles the Python surface, runs all tests, checks Quarto/Markdown consistency, verifies the published bundle and artifacts, checks local documentation links, replays the primary conformal analysis, and compares it with the tracked tables.
  • The exact local release gate is:
quarto render docs/safe_openvla_audit.qmd --to gfm
uv run --frozen --extra dev python scripts/build_publication_manifest.py
uv run --frozen --extra dev python -m compileall -q data models scripts tests
uv run --frozen --extra dev pytest -q
uv run --frozen --extra dev python scripts/verify_publication.py

Rebuild the manifest only after intentional edits; CI verifies the committed manifest without regenerating it. The score replay and table comparison in DATA.md are also required before publication.