# Native SmolVLA Model Migration Design ## Goal Migrate only the SmolVLA model core from `/data/lerobot-imf-attnres-exp/lerobot-imf-attnres` into this `roboimi` project so Diana simulation can train/evaluate through the existing `roboimi.vla` agent interface without importing or installing the full LeRobot tree. ## Non-goals - Do not vendor the full `lerobot` package. - Do not change the current Python environment to LeRobot 0.5.x requirements. - Do not alter Diana environment action semantics. - Do not change `train_vla.py` or `eval_vla.py` main control flow unless a minimal compatibility hook is unavoidable. - Do not implement RTC in the first pass. ## Architecture Add a native RoboIMI SmolVLA package under `roboimi/vla/models/smolvla/`. It will contain a lightweight config dataclass, a minimally adapted `SmolVLMWithExpertModel`, and a `VLAFlowMatching` implementation. A new `SmolVLANativeAgent` wraps the model with the existing RoboIMI agent contract: `compute_loss`, `predict_action_chunk`, `select_action`, `reset`, and `get_normalization_stats`. The new code will preserve the original model math where practical, but replace LeRobot framework dependencies with local constants, tokenizer handling, queue management, and `roboimi.vla.models.normalization.NormalizationModule`. ## Data flow Training batch input remains the current RoboIMI format: ```python { "images": {cam: Tensor[B, T, C, H, W]}, "qpos": Tensor[B, T, 16], "action": Tensor[B, H, 16], "action_is_pad": optional BoolTensor[B, H], "task": optional list[str], } ``` `SmolVLANativeAgent` normalizes `qpos` and `action` using RoboIMI dataset stats, tokenizes task strings, and passes images/state/language/action into the native SmolVLA model. In inference, the model returns normalized action chunks; the agent denormalizes them to 16-dim Diana EE actions. ## Components ### `roboimi/vla/models/smolvla/configuration.py` Defines `NativeSmolVLAConfig`, a small dataclass with fields needed by the model: state/action dimensions, padding dimensions, image resize target, tokenizer settings, VLM model name, VLM loading flags, expert layer configuration, sampling steps, dtype/device behavior, and compile flags. ### `roboimi/vla/models/smolvla/smolvlm_with_expert.py` Migrates the model-core helper from the old repo. It should depend only on PyTorch and Transformers. It will expose `SmolVLMWithExpertModel` and attention helpers. ### `roboimi/vla/models/smolvla/modeling.py` Defines model-core utilities (`resize_with_pad`, `pad_vector`, `make_att_2d_masks`, sinusoidal time embedding) plus `VLAFlowMatching`, with `forward` and `sample_actions`. ### `roboimi/vla/agent_smolvla_native.py` RoboIMI-native agent wrapper. It owns tokenizer, normalization, task fallback, camera ordering, observation/action queues, loss masking, and denormalized rollout actions. ### `roboimi/vla/conf/agent/smolvla_native.yaml` Hydra config for the native model, defaulting to Diana 16-dim state/action and configurable camera names. ## Error handling - Missing configured camera raises `ValueError` with expected/missing names. - Task list length not matching batch size raises `ValueError`. - State/action dimensions exceeding configured `max_state_dim` / `max_action_dim` raises `ValueError`. - Missing Transformers SmolVLM classes raises `ImportError` explaining the required package. - Invalid action chunk shape raises `RuntimeError` explaining expected `(B,H,A)`. ## Testing Use TDD with fake VLM/tokenizer/model components first, so tests do not download or instantiate the real SmolVLM. Cover: 1. Config and Hydra instantiation with fake injected components. 2. Camera ordering and missing-camera errors. 3. Task fallback and newline/tokenization behavior. 4. Loss path shape/mask behavior using a fake native model. 5. `predict_action_chunk` normalization/denormalization and shape. 6. `select_action` queue behavior. A later smoke test may instantiate the real model in an environment that already has the required Transformers/weights, but the first implementation must pass without external downloads. ## Acceptance criteria - `agent=smolvla_native` can be composed by Hydra. - Unit tests pass without importing external `/data/.../src/lerobot`. - The new production code has no `import lerobot`. - The agent accepts the same batch/observation structure used by current train/eval scripts. - The agent emits 16-dim denormalized Diana EE actions for rollout.