What happened

Google DeepMind introduced the Gemini Robotics 2 family on July 30. The release divides the system into three models rather than presenting one model as the entire robotics stack.

Gemini Robotics 2 is a vision-language-action model for whole-body control. Gemini Robotics ER 2 handles physical-world understanding, communication, multi-step planning, and multi-robot coordination. Gemini Robotics On-Device 2 is optimized for local execution on robot hardware.

From demonstrations to a usable stack

Google showed whole-body walking and manipulation on Apptronik’s Apollo 2 humanoid, along with dexterous tasks on other embodiments. The company also acknowledged that movement speed still needs improvement.

The access model reinforces the experimental status. ER 2 is available in Google AI Studio and in private preview on the Gemini Enterprise Agent Platform. The whole-body VLA and on-device model are available only to early-access partners.

Why it matters

Separating orchestration, action, and local inference reflects the practical architecture of a robot. High-level planning and human interaction have different latency and compute needs from continuous motor control.

Google says On-Device 2 can adapt to new bi-arm embodiments using a few hours of data and typically fewer than 200 examples. That result is a company claim; wider access and independent evaluation will be needed to understand how it holds across hardware and tasks.

What to watch next

  • Movement speed and reliability outside curated demonstrations.
  • A broader release path for the whole-body and on-device models.
  • Independent results for adaptation across unfamiliar robot bodies.
  • Multi-robot coordination in sustained physical workflows.