203 lines
5.8 KiB
Markdown
203 lines
5.8 KiB
Markdown
# VLA Experiment Sweep Execution Plan
|
|
|
|
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
|
|
|
|
**Goal:** Probe the largest safe shared batch size, prepare the 10 unique VLA sweep runs, and launch the serial GPU experiment queue with SwanLab logging and 20-epoch headless rollout validation.
|
|
|
|
**Architecture:** Use the existing `train_vla.py` / `eval_vla.py` path without changing model code unless launch-blocking issues appear. A small external launcher contract will manage per-run directories, logs, pids, and serial execution while Hydra places trainer-local `checkpoints/` under each run directory.
|
|
|
|
**Tech Stack:** zsh, mamba env `roboimi`, Hydra overrides, PyTorch CUDA, SwanLab, existing `train_vla.py`
|
|
|
|
---
|
|
|
|
### Task 1: Reconfirm schedule math and current runtime assumptions
|
|
|
|
**Files:**
|
|
- Modify: `docs/superpowers/specs/2026-03-30-vla-experiment-sweep-design.md` (only if execution reveals a spec mismatch)
|
|
- Test: none
|
|
|
|
- [ ] **Step 1: Print dataset size and epoch math for the baseline batch size**
|
|
|
|
Run:
|
|
```bash
|
|
python - <<'PY'
|
|
num_samples = 70000
|
|
for batch_size in (32, 48, 64, 80):
|
|
steps_per_epoch = num_samples // batch_size
|
|
print(batch_size, steps_per_epoch, 20 * steps_per_epoch, 200 * steps_per_epoch)
|
|
PY
|
|
```
|
|
|
|
Expected: integer epoch-step mappings for candidate batch sizes.
|
|
|
|
- [ ] **Step 2: Reconfirm no stale training process is already active**
|
|
|
|
Run:
|
|
```bash
|
|
ps -ef | grep -E 'train_vla.py|python .*roboimi/demos/vla_scripts/train_vla.py' | grep -v grep || true
|
|
```
|
|
|
|
Expected: no active training process unless intentionally started by this session.
|
|
|
|
- [ ] **Step 3: Record the execution assumptions**
|
|
|
|
Capture in session notes:
|
|
- dataset path
|
|
- camera list
|
|
- current branch
|
|
- no active conflicting training process
|
|
|
|
|
|
### Task 2: Probe the largest safe shared batch size
|
|
|
|
**Files:**
|
|
- Create: `runs/batch-probe-<timestamp>/` (runtime artifact)
|
|
- Test: probe command output only
|
|
|
|
- [ ] **Step 1: Write the candidate probe command for the largest model**
|
|
|
|
Use overrides:
|
|
- `data.dataset_dir=/home/droid/project/diana_sim/sim_transfer`
|
|
- `data.camera_names=[r_vis,top,front]`
|
|
- `agent.head.n_emb=384`
|
|
- `agent.head.n_layer=12`
|
|
- `agent.head.n_head=4`
|
|
- `agent.vision_backbone.pretrained_backbone_weights=IMAGENET1K_V1`
|
|
- `agent.vision_backbone.freeze_backbone=true`
|
|
- `agent.vision_backbone.use_separate_rgb_encoder_per_camera=true`
|
|
- `train.device=cuda`
|
|
- `train.val_split=0.0`
|
|
- `train.seed=42`
|
|
- `train.use_swanlab=false`
|
|
- `train.rollout_val_freq_epochs=0`
|
|
- `train.rollout_validate_on_checkpoint=false`
|
|
- `train.max_steps=4`
|
|
|
|
- [ ] **Step 2: Run the probe for `batch_size=32`**
|
|
|
|
Expected pass condition:
|
|
- no CUDA OOM
|
|
- no crash/abort
|
|
- no NaN/Inf loss
|
|
- 4 steps complete
|
|
|
|
- [ ] **Step 3: Repeat for `batch_size=48`, `64`, then optionally `80`**
|
|
|
|
Stop increasing once one candidate fails.
|
|
|
|
- [ ] **Step 4: Compute the shared LR and 200-epoch step budget**
|
|
|
|
Using:
|
|
```text
|
|
lr = 1e-4 * (batch_size / 16)
|
|
max_steps = 200 * floor(70000 / batch_size)
|
|
```
|
|
|
|
- [ ] **Step 5: Write down the chosen shared batch size and derived max_steps**
|
|
|
|
This becomes the execution contract for all runs.
|
|
|
|
|
|
### Task 3: Materialize the 10 unique run matrix
|
|
|
|
**Files:**
|
|
- Create: `runs/vla-sweep-<timestamp>/manifest.txt` (runtime artifact)
|
|
- Create: `runs/vla-sweep-<timestamp>/launch_queue.sh` (runtime artifact)
|
|
- Test: shell syntax / manifest inspection
|
|
|
|
- [ ] **Step 1: Enumerate the 2-point pretraining comparison**
|
|
|
|
Canonical runs:
|
|
1. `sim-transfer-baseline-pretrain-on-emb128-layer4`
|
|
2. `sim-transfer-pretrain-off`
|
|
|
|
- [ ] **Step 2: Enumerate the remaining 8 unique architecture-sweep runs**
|
|
|
|
Skip the duplicate `(pretrain-on, emb128, layer4)` because it reuses the baseline run.
|
|
|
|
- [ ] **Step 3: For each run, define exact Hydra overrides**
|
|
|
|
Every run must pin:
|
|
- dataset path
|
|
- three cameras
|
|
- `train.device=cuda`
|
|
- `train.val_split=0.0`
|
|
- `train.num_workers=12`
|
|
- `train.seed=42`
|
|
- `train.rollout_val_freq_epochs=20`
|
|
- `train.rollout_validate_on_checkpoint=false`
|
|
- `train.rollout_num_episodes=3`
|
|
- `train.use_swanlab=true`
|
|
- a unique `train.swanlab_run_name`
|
|
- shared batch size
|
|
- shared LR
|
|
- derived `max_steps`
|
|
- `agent.head.n_head=4`
|
|
- `agent.vision_backbone.freeze_backbone=true`
|
|
- `agent.vision_backbone.use_separate_rgb_encoder_per_camera=true`
|
|
|
|
- [ ] **Step 4: Write the queue launcher script**
|
|
|
|
Launcher responsibilities:
|
|
- create run directory
|
|
- export `LD_LIBRARY_PATH` CUDA/cuDNN fixup
|
|
- export `MPLCONFIGDIR=/tmp/mpl`
|
|
- set `hydra.job.chdir=true`
|
|
- set `hydra.run.dir=<run_dir>`
|
|
- tee logs to `<run_dir>/train.log`
|
|
- write `<run_dir>/train.pid`
|
|
- run jobs serially
|
|
|
|
- [ ] **Step 5: Verify the queue script syntax**
|
|
|
|
Run:
|
|
```bash
|
|
zsh -n runs/vla-sweep-<timestamp>/launch_queue.sh
|
|
```
|
|
|
|
Expected: exit 0.
|
|
|
|
|
|
### Task 4: Launch the first probe-selected production run and verify startup
|
|
|
|
**Files:**
|
|
- Create: `runs/vla-sweep-<timestamp>/<run-name>/` (runtime artifact)
|
|
- Test: live log inspection / process inspection
|
|
|
|
- [ ] **Step 1: Start the first run in the queue**
|
|
|
|
Use the generated launcher contract.
|
|
|
|
- [ ] **Step 2: Inspect the first 100-200 log lines**
|
|
|
|
Expected:
|
|
- config resolves correctly
|
|
- dataset loads
|
|
- agent initializes on CUDA
|
|
- training loop starts
|
|
- SwanLab initializes
|
|
|
|
- [ ] **Step 3: Confirm process identity and run directory**
|
|
|
|
Collect:
|
|
- PID
|
|
- run dir
|
|
- log path
|
|
- SwanLab run name
|
|
|
|
- [ ] **Step 4: If startup is clean, continue the remaining serial queue**
|
|
|
|
If startup fails, stop and debug before launching more runs.
|
|
|
|
|
|
### Task 5: Report the launched sweep contract back to the user
|
|
|
|
**Files:**
|
|
- Test: log evidence only
|
|
|
|
- [ ] **Step 1: Summarize the chosen batch size, LR, max_steps, and rollout cadence**
|
|
|
|
- [ ] **Step 2: Provide the active run directory / pid / first-run status**
|
|
|
|
- [ ] **Step 3: State the remaining queued runs by name**
|