VLA models

SmolVLA

United States

A 450M open VLA trained only on crowd-sourced community data, small enough to train on a single consumer GPU and run on CPU or a MacBook; asynchronous inference yields ~30% faster response and ~2x task throughput. Built by Hugging Face's Paris-based LeRobot team.

Updated 2026-07-09 · Last verified 2026-07-09

OrganizationHugging Face (LeRobot team)
CountryUS
Release2025-06
Params (B)0.5
Open weightsyes
Architecturecompact VLA: SmolVLM-2 backbone + flow-matching action expert conditioned on multi-camera views, robot state and language; asynchronous inference
LicenseApache-2.0
Modalitiesvision, language, action
Embodimentsmanipulator
Training dataExclusively publicly available, crowd-sourced LeRobot community datasets from the Hugging Face hub (low-cost robot data)

Sources : arXiv (2025-06-02)Hugging Face (2025-06-03)Hugging Face (model card) (2026-07-09)

← Back to comparator · VLA models
Christian Verbrugge

Christian Verbrugge · D·Fairy

Bringing a physical AI solution into European industry?

15 years at KUKA on automotive OEM programs (€100M), now focused on edge AI and autonomous agents. I support robotics and embodied AI companies on their EMEA market entry: OEM and Tier 1 access, co-funded pilots, AI Act readiness.

Book an intro call