VLA models

RT-2 (Robotic Transformer 2)

United States

The founding VLA: first to show that web-scale vision-language knowledge transfers directly to robot control by treating actions as text tokens. It spawned the RT-X / Open X-Embodiment cross-lab effort that underpins most later open VLAs. Research artifact, never released as weights or product.

Updated 2026-07-09 · Last verified 2026-07-09

OrganizationGoogle DeepMind
CountryUS
Release2023-07
Params (B)55
Open weightsno
ArchitectureVLA co-fine-tuned from web-scale VLMs (PaLI-X 55B and PaLM-E 12B variants); robot actions emitted as text tokens
Modalitiesvision, language, action
Embodimentsmanipulator, mobile
Training dataInternet-scale vision-language data co-fine-tuned with RT-1 robot demonstrations (collected with 13 robots over 17 months)

Sources : Google DeepMind (2023-07-28)arXiv (2023-07-28)InfoQ (2023-10-01)

← Back to comparator · VLA models
Christian Verbrugge

Christian Verbrugge · D·Fairy

Bringing a physical AI solution into European industry?

15 years at KUKA on automotive OEM programs (€100M), now focused on edge AI and autonomous agents. I support robotics and embodied AI companies on their EMEA market entry: OEM and Tier 1 access, co-funded pilots, AI Act readiness.

Book an intro call