Files
roboimi/workspace/inputs/experimental_log.md
T

11 KiB

Experimental Log: iMF-AttnRes for Fast Vision-Language-Action Imitation Learning in RoboIMI

1. Experimental Setup

We evaluate imitation-learning policies in the RoboIMI simulated manipulation benchmark. The study focuses on two environments:

  • socket_peg / socket-insert: a socket peg insertion task.
  • sim_transfer / object transfer: a block/object transfer manipulation task.

The central method is an iMF-AttnRes VLA policy that combines improved Mean Flow (iMF) training with Attention Residuals (AttnRes). The comparison set includes diffusion-policy-style DiT baselines, ACT, SmolVLA, and several ablations of the iMF-AttnRes architecture and horizon/execution settings. Each reported result is based on 100 simulation rollouts unless otherwise stated.

Reported metrics:

  • avg_reward: mean cumulative episode reward over 100 rollouts.
  • median_reward: median cumulative episode reward.
  • avg_max_reward: mean maximum per-episode reward statistic reported by the rollout evaluator.
  • max_reward: highest cumulative episode reward among 100 rollouts.
  • nonzero_reward: number of episodes with nonzero reward.
  • success_like: number of episodes whose max reward crosses the environment threshold (max_reward > 4 for socket-insert; max_reward >= 4 for sim_transfer, as recorded in the logs).
  • avg_inference_fps, avg_control_fps, and avg_inference_time_ms: rollout speed measurements. Some sim_transfer logs report FPS but not inference time in milliseconds.

Hardware varies across runs and is encoded in run names: 5880g0/5880g1 denotes RTX 5880 Ada runs, l20g* denotes L20 runs, and 5090 denotes RTX 5090 runs. Because hardware differs across experiments, speed comparisons should be interpreted as practical rollout measurements rather than perfectly hardware-normalized measurements.

2. Raw Numeric Data

Table 1: Socket peg insertion, 100-rollout performance and speed

run_name method_family horizon_or_steps avg_reward median_reward avg_max_reward max_reward nonzero_reward success_like avg_inference_fps avg_control_fps avg_inference_time_ms
socket-insert-imf-attnres-ph32-exec16-emb384-l12-infer1-unfreeze-step50k-roll1x10-5880g01-20260506-195806 iMF-AttnRes infer1, ph32, exec16, 50k 1275.22 1490.5 3.12 2250.0 84/100 0/100 64.226 5.863 16.186
socket-insert-no-pretrain-ph16-emb384-l18-infer100-unfreeze-step150k-roll1x10-5090-20260506-193753 Diffusion-style VLA infer100, ph16, 150k 861.53 668.0 2.58 2241.0 97/100 2/100 2.956 2.591 338.291
socket-insert-no-pretrain-ph32-exec16-emb384-l18-infer100-unfreeze-step150k-roll1x10-5880g1-20260507-092227 Diffusion-style VLA infer100, ph32, exec16, 150k 1019.39 804.0 2.47 2385.0 92/100 4/100 2.516 1.928 397.912
socket-insert-imf-attnres-ph32-exec16-emb384-l12-infer2-drop005-unfreeze-step150k-roll1x10-5880g01-20260508-165022 iMF-AttnRes infer2, ph32, exec16, 150k 1472.62 1825.0 3.29 2365.0 83/100 5/100 71.579 7.181 14.409
socket-insert-imf-attnres-ph32-exec16-emb384-l12-infer3-drop005-unfreeze-step150k-roll1x10-5880g01-20260509-172423 iMF-AttnRes infer3, ph32, exec16, 150k 1513.56 1901.5 3.28 2285.0 83/100 3/100 67.891 7.247 15.122
act-socket-peg-224-20260508-170237 ACT action chunking 289.6 13.0 1.29 2091.0 55/100 1/100 10.5171 2.2782 100.2116
smolvla_socket_peg_bs80_100k_20260511_145727 SmolVLA 100k 466.16 106.0 1.72 4.0 89/100 0/100 382.53 16.84 2.614

Table 2: Object transfer / sim_transfer, 100-rollout performance and speed

run_name method_family horizon_or_steps avg_reward median_reward avg_max_reward max_reward nonzero_reward success_like avg_inference_fps avg_control_fps avg_inference_time_ms
diffusion_policy_native_dit_ddpm_resnet_best Diffusion Policy native DiT + DDPM + ResNet best checkpoint 319.2 6.0 1.68 1416.0 55/100 29/100 32.09 16.70 n/a
embed384_layer18_best_checkpoint Diffusion-style VLA emb384, layer18 233.52 0.0 1.12 1436.0 39/100 17/100 1.859 1.649 n/a
resnet18-multitoken-imf-emb256-l16-ph16-ex08-roll10-l20g2-20260406-112815 iMF multi-token ResNet18 step34999, ph16, exec08 260.66 0.0 1.12 1422.0 33/100 23/100 441.80 13.37 n/a
imf-p2-full-attnres-vision-ph16-ex08-emb384-l12-b40-lr1p25e4-ms50k-l20g3-20260405-002424 iMF full AttnRes vision ph16, exec08, 50k 228.42 0.0 0.94 1482.0 31/100 16/100 55.997 4.301 n/a
imf-p1-ph16-ex16-emb384-l12-ms50k-l20g0-20260404-131223 iMF-AttnRes DiT only ph16, exec16, 50k 240.64 0.0 1.02 1428.0 32/100 19/100 137.996 4.523 n/a
imf-p1-ph32-ex08-emb384-l12-ms50k-l20g1-20260404-131223 iMF-AttnRes DiT only ph32, exec08, 50k 163.28 0.0 0.76 1566.0 26/100 12/100 85.638 4.317 n/a
imf-p1-ph32-ex32-emb384-l12-ms50k-5090-20260404-13122 iMF-AttnRes DiT only ph32, exec32, 50k 260.72 0.0 1.18 1342.0 38/100 21/100 9.557 3.299 n/a
imf-p1-ph16-ex08-emb384-l12-ms50k-5880g1-20260404-131223 iMF-AttnRes DiT only ph16, exec08, 50k 229.56 0.0 1.18 1564.0 41/100 18/100 69.984 8.053 n/a
imf-p1-ph08-ex08-emb384-l12-ms50k-5880g0-20260404-131223 iMF-AttnRes DiT only ph08, exec08, 50k 237.02 0.0 1.32 1400.0 48/100 18/100 69.192 8.173 n/a
imf-p1-ph32-ex16-emb384-l12-ms50k-l20g2-20260404-131223 iMF-AttnRes DiT only ph32, exec16, 50k 49.88 0.0 0.32 1284.0 13/100 4/100 138.903 4.482 n/a
sim-transfer-imf-attnres-ph32-exec16-emb384-l12-infer1-unfreeze-step50k-roll5x5-20260403-14130 iMF-AttnRes infer1, ph32, exec16, 50k 526.22 86.0 2.14 1460.0 63/100 44/100 8.911 3.193 n/a

Table 3: Derived comparisons used in the paper draft

comparison metric numerator_run denominator_run numerator_value denominator_value derived_ratio_or_delta
socket best iMF vs ph16 diffusion baseline avg_reward ratio iMF infer3 150k diffusion ph16 infer100 150k 1513.56 861.53 1.76x
socket best iMF vs ph32 diffusion baseline avg_reward ratio iMF infer3 150k diffusion ph32 infer100 150k 1513.56 1019.39 1.48x
socket best iMF vs ACT avg_reward ratio iMF infer3 150k ACT 1513.56 289.6 5.23x
socket best iMF vs SmolVLA avg_reward ratio iMF infer3 150k SmolVLA 1513.56 466.16 3.25x
socket iMF infer2 vs ph16 diffusion baseline inference time reduction iMF infer2 150k diffusion ph16 infer100 150k 14.409 338.291 23.48x lower latency
socket iMF infer3 vs ph16 diffusion baseline inference fps ratio iMF infer3 150k diffusion ph16 infer100 150k 67.891 2.956 22.97x
socket iMF infer2 vs ph32 diffusion baseline inference fps ratio iMF infer2 150k diffusion ph32 infer100 150k 71.579 2.516 28.45x
sim_transfer best iMF vs native diffusion policy avg_reward ratio sim-transfer iMF infer1 native diffusion policy 526.22 319.2 1.65x
sim_transfer best iMF vs native diffusion policy success_like delta sim-transfer iMF infer1 native diffusion policy 44 29 +15 episodes
sim_transfer ResNet18 multitoken iMF vs embed384 layer18 baseline inference fps ratio ResNet18 multitoken iMF embed384 layer18 441.80 1.859 237.65x

3. Qualitative Observations

Socket peg insertion shows the clearest quality-speed tradeoff improvement. The iMF-AttnRes variants with ph32/exec16 and only 1--3 inference evaluations achieve substantially higher average and median rewards than the diffusion-style infer100 baselines, while reducing latency from hundreds of milliseconds to roughly 14--16 ms. The infer2 and infer3 variants are the strongest socket policies by average reward, with the infer3 run obtaining avg_reward 1513.56 and median_reward 1901.5. The infer2 run has the fastest iMF-AttnRes socket latency among the high-reward variants at 14.409 ms.

The socket experiments also reveal that nonzero reward alone is not a sufficient quality metric. The diffusion-style ph16 baseline has 97/100 nonzero episodes but much lower median reward than the best iMF-AttnRes runs. Conversely, iMF-AttnRes has 83/100 nonzero episodes but much higher median and average reward, suggesting that when it engages successfully with the task it produces stronger progress toward insertion.

SmolVLA is extremely fast in inference FPS and control FPS, but its reward is lower than iMF-AttnRes in these experiments. ACT is slower than SmolVLA and underperforms iMF-AttnRes in average and median reward.

In sim_transfer, the best recorded iMF-AttnRes run, sim-transfer-imf-attnres-ph32-exec16-emb384-l12-infer1-unfreeze-step50k-roll5x5-20260403-14130, reaches avg_reward 526.22 and 44/100 success-like episodes, exceeding the native diffusion-policy DiT/DDPM/ResNet baseline with avg_reward 319.2 and 29/100 success-like episodes. However, several iMF ablations underperform the native diffusion-policy baseline on average reward, indicating that the method is sensitive to horizon/execution choices, vision-token design, and whether AttnRes is applied only in the policy transformer or throughout the vision stack.

For sim_transfer, the ResNet18 multi-token iMF variant is a particularly fast policy at 441.80 inference FPS but does not match the best iMF-AttnRes reward. The full-AttnRes vision variant underperforms the more conservative DiT-only AttnRes settings, suggesting that replacing residual pathways inside the vision encoder may be harmful or require different training hyperparameters.

Iteration History

  1. Early sim_transfer experiments compared diffusion-style baselines against initial iMF-AttnRes policy variants and explored policy horizon/execution length combinations.
  2. Subsequent ablations tested multi-token ResNet18 visual encoding, full AttnRes in the vision backbone, and DiT-only AttnRes. These results showed that iMF speedups are robust but reward quality is sensitive to model and horizon choices.
  3. Later socket-insert experiments scaled to 100-rollout evaluation and compared iMF-AttnRes against diffusion-style VLA, ACT, and SmolVLA. These experiments provided the strongest evidence for the paper's main claim: iMF-AttnRes can substantially reduce inference latency while improving or preserving manipulation reward.

4. Caveats for Paper Writing

  • Do not claim real-robot validation; all reported experiments are simulation rollouts.
  • Hardware differs across runs, so speed measurements are useful but not fully hardware-normalized.
  • Success-like episode counts use thresholds as recorded by each environment's evaluation script: socket-insert logs report max_reward > 4, whereas sim_transfer logs report max_reward >= 4.
  • The reported data are sufficient for a first paper draft but should be supplemented later with controlled hardware-normalized speed measurements and more seeds if the paper is prepared for formal submission.