# VLA Experiment Sweep Execution Plan > **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. **Goal:** Probe the largest safe shared batch size, prepare the 10 unique VLA sweep runs, and launch the serial GPU experiment queue with SwanLab logging and 20-epoch headless rollout validation. **Architecture:** Use the existing `train_vla.py` / `eval_vla.py` path without changing model code unless launch-blocking issues appear. A small external launcher contract will manage per-run directories, logs, pids, and serial execution while Hydra places trainer-local `checkpoints/` under each run directory. **Tech Stack:** zsh, mamba env `roboimi`, Hydra overrides, PyTorch CUDA, SwanLab, existing `train_vla.py` --- ### Task 1: Reconfirm schedule math and current runtime assumptions **Files:** - Modify: `docs/superpowers/specs/2026-03-30-vla-experiment-sweep-design.md` (only if execution reveals a spec mismatch) - Test: none - [ ] **Step 1: Print dataset size and epoch math for the baseline batch size** Run: ```bash python - <<'PY' num_samples = 70000 for batch_size in (32, 48, 64, 80): steps_per_epoch = num_samples // batch_size print(batch_size, steps_per_epoch, 20 * steps_per_epoch, 200 * steps_per_epoch) PY ``` Expected: integer epoch-step mappings for candidate batch sizes. - [ ] **Step 2: Reconfirm no stale training process is already active** Run: ```bash ps -ef | grep -E 'train_vla.py|python .*roboimi/demos/vla_scripts/train_vla.py' | grep -v grep || true ``` Expected: no active training process unless intentionally started by this session. - [ ] **Step 3: Record the execution assumptions** Capture in session notes: - dataset path - camera list - current branch - no active conflicting training process ### Task 2: Probe the largest safe shared batch size **Files:** - Create: `runs/batch-probe-/` (runtime artifact) - Test: probe command output only - [ ] **Step 1: Write the candidate probe command for the largest model** Use overrides: - `data.dataset_dir=/home/droid/project/diana_sim/sim_transfer` - `data.camera_names=[r_vis,top,front]` - `agent.head.n_emb=384` - `agent.head.n_layer=12` - `agent.head.n_head=4` - `agent.vision_backbone.pretrained_backbone_weights=IMAGENET1K_V1` - `agent.vision_backbone.freeze_backbone=true` - `agent.vision_backbone.use_separate_rgb_encoder_per_camera=true` - `train.device=cuda` - `train.val_split=0.0` - `train.seed=42` - `train.use_swanlab=false` - `train.rollout_val_freq_epochs=0` - `train.rollout_validate_on_checkpoint=false` - `train.max_steps=4` - [ ] **Step 2: Run the probe for `batch_size=32`** Expected pass condition: - no CUDA OOM - no crash/abort - no NaN/Inf loss - 4 steps complete - [ ] **Step 3: Repeat for `batch_size=48`, `64`, then optionally `80`** Stop increasing once one candidate fails. - [ ] **Step 4: Compute the shared LR and 200-epoch step budget** Using: ```text lr = 1e-4 * (batch_size / 16) max_steps = 200 * floor(70000 / batch_size) ``` - [ ] **Step 5: Write down the chosen shared batch size and derived max_steps** This becomes the execution contract for all runs. ### Task 3: Materialize the 10 unique run matrix **Files:** - Create: `runs/vla-sweep-/manifest.txt` (runtime artifact) - Create: `runs/vla-sweep-/launch_queue.sh` (runtime artifact) - Test: shell syntax / manifest inspection - [ ] **Step 1: Enumerate the 2-point pretraining comparison** Canonical runs: 1. `sim-transfer-baseline-pretrain-on-emb128-layer4` 2. `sim-transfer-pretrain-off` - [ ] **Step 2: Enumerate the remaining 8 unique architecture-sweep runs** Skip the duplicate `(pretrain-on, emb128, layer4)` because it reuses the baseline run. - [ ] **Step 3: For each run, define exact Hydra overrides** Every run must pin: - dataset path - three cameras - `train.device=cuda` - `train.val_split=0.0` - `train.num_workers=12` - `train.seed=42` - `train.rollout_val_freq_epochs=20` - `train.rollout_validate_on_checkpoint=false` - `train.rollout_num_episodes=3` - `train.use_swanlab=true` - a unique `train.swanlab_run_name` - shared batch size - shared LR - derived `max_steps` - `agent.head.n_head=4` - `agent.vision_backbone.freeze_backbone=true` - `agent.vision_backbone.use_separate_rgb_encoder_per_camera=true` - [ ] **Step 4: Write the queue launcher script** Launcher responsibilities: - create run directory - export `LD_LIBRARY_PATH` CUDA/cuDNN fixup - export `MPLCONFIGDIR=/tmp/mpl` - set `hydra.job.chdir=true` - set `hydra.run.dir=` - tee logs to `/train.log` - write `/train.pid` - run jobs serially - [ ] **Step 5: Verify the queue script syntax** Run: ```bash zsh -n runs/vla-sweep-/launch_queue.sh ``` Expected: exit 0. ### Task 4: Launch the first probe-selected production run and verify startup **Files:** - Create: `runs/vla-sweep-//` (runtime artifact) - Test: live log inspection / process inspection - [ ] **Step 1: Start the first run in the queue** Use the generated launcher contract. - [ ] **Step 2: Inspect the first 100-200 log lines** Expected: - config resolves correctly - dataset loads - agent initializes on CUDA - training loop starts - SwanLab initializes - [ ] **Step 3: Confirm process identity and run directory** Collect: - PID - run dir - log path - SwanLab run name - [ ] **Step 4: If startup is clean, continue the remaining serial queue** If startup fails, stop and debug before launching more runs. ### Task 5: Report the launched sweep contract back to the user **Files:** - Test: log evidence only - [ ] **Step 1: Summarize the chosen batch size, LR, max_steps, and rollout cadence** - [ ] **Step 2: Provide the active run directory / pid / first-run status** - [ ] **Step 3: State the remaining queued runs by name**