VLA Models
VLA (Vision-Language-Action) – “VLM + action output”: see and reason plus act with hands.
Two-level architecture:
- slow long-term planning – what to do: decompose instruction into subgoals;
- fast VLA control – how to do: continuous control signals.
Two directions of action representation:
- discrete tokens (RT-2, OpenVLA) – autoregression as text generation;
- continuous trajectory (π₀) – flow matching, smoother for precise manipulations.
Foundation – imitation learning on cross-platform demonstrations; RL – an additional layer. Action decomposition into blocks: one model output – a sequence of future actions, execution at high frequency.
Related: Computer Use, Sim2Real, Two-level planning-control architecture