VLA models

AgiBot GO-1 (Genie Operator-1)

China

China's flagship open embodied foundation model, introducing the ViLLA framework where latent action tokens bridge vision-language inputs and robot control. Announced March 2025 and fully open-sourced (weights on Hugging Face) in September 2025, alongside the largest open real-robot dataset.

Updated 2026-07-09 · Last verified 2026-07-09

OrganizationAgiBot (Zhiyuan Robotics) / OpenDriveLab
CountryCN
Release2025-03
Params (B)3
Open weightsyes
ArchitectureViLLA (Vision-Language-Latent-Action): InternVL2.5-2B VLM backbone + MoE with a 24-layer latent planner predicting latent action tokens and a high-frequency action expert
LicenseCC BY-NC-SA 4.0
Modalitiesvision, language, action
Embodimentshumanoid, manipulator, cross-embodiment
Training dataAgiBot World dataset: over 1 million real-robot trajectories across 217 tasks in 5 application domains, plus cross-embodiment and human video data for the latent planner

Sources : Hugging Face (model card) (2026-07-09)GlobeNewswire (AgiBot press release) (2025-03-11)arXiv (2025-03-09)OpenDriveLab (X) (2025-09-20)

← Back to comparator · VLA models
Christian Verbrugge

Christian Verbrugge · D·Fairy

Bringing a physical AI solution into European industry?

15 years at KUKA on automotive OEM programs (€100M), now focused on edge AI and autonomous agents. I support robotics and embodied AI companies on their EMEA market entry: OEM and Tier 1 access, co-funded pilots, AI Act readiness.

Book an intro call