VLA models
RT-2 (Robotic Transformer 2)
The founding VLA: first to show that web-scale vision-language knowledge transfers directly to robot control by treating actions as text tokens. It spawned the RT-X / Open X-Embodiment cross-lab effort that underpins most later open VLAs. Research artifact, never released as weights or product.
Sources : Google DeepMind (2023-07-28)arXiv (2023-07-28)InfoQ (2023-10-01)
← Back to comparator · VLA models