Gemini Robotics 2: Why "Whole-Body Control" Is the Actual News, Not the Humanoid Hype
2026 has produced a steady stream of humanoid robot demo videos, and it's easy to get numb to them: a robot folds a shirt, a robot walks across a stage, a robot runs 100m in 8.86 seconds at the World Humanoid Robot Games in Beijing. Google DeepMind's Gemini Robotics 2, announced July 30, 2026, is easy to lump into that same pile of demo footage. Underneath the footage is an architecture change, and "whole-body control" names a problem humanoid labs have been stuck on for years.
Upper bodies on fixed bases
Most of the humanoid demos that made headlines through 2025 and early 2026 (including earlier Gemini Robotics models, which by DeepMind's own description "controlled the humanoid's upper-body to achieve table-top tasks") controlled a robot's upper body: arms, hands, sometimes a torso, while the legs either didn't exist, were bolted to a fixed base, or ran on a separate, much simpler locomotion controller that had no idea what the arms were doing. That split explains why a robot could look dexterous in a tabletop demo and immediately look absurd the moment it had to bend down, brace against something, or use its whole body's momentum to lift or push, because the system that decided how to move the legs and the system that decided how to move the arms weren't coordinated.
Gemini Robotics 2 is Google DeepMind's attempt to close that gap: one system that controls a humanoid "from feet to fingertips," coordinating walking and crouching with five-finger dexterity (the announcement shows Apptronik's Apollo 2 doing this, plus delicate hand tasks like tying knots or sealing a ziplock bag).
Three models behind one name
DeepMind released three models under the Gemini Robotics 2 name, and each has a different job:
- Gemini Robotics 2 (the VLA): a vision-language-action model that converts vision and language input into motor control, for full humanoids as well as bi-arm robots with dexterous manipulation.
- Gemini Robotics ER 2: the embodied-reasoning model DeepMind calls "the robot's high-level brain." It processes user instructions, reasons about the steps needed, and coordinates with the VLA to carry them out, supporting tasks that last several minutes and involve hundreds of decisions.
- Gemini Robotics On-Device 2: DeepMind's most efficient VLA, optimized to run locally on the robot for cases without reliable connectivity or with tight latency needs, and adaptable to a new robot body in a few hours, typically with fewer than 200 examples.
Splitting reasoning (ER 2) from low-level motor control (the VLA) is the same pattern that's shown up repeatedly in robotics ML over the past few years: a slow, deliberate planner feeding a fast, reactive controller, because a single model trying to do both tends to be bad at either the long-horizon reasoning or the millisecond-scale control loop. What's new here is doing it across a full humanoid body, with a third, efficiency-focused variant built to run on the robot itself. DeepMind doesn't say what the smaller model gives up. Smaller on-robot models usually trade away some capability, so published comparisons against the full VLA would be the thing to wait for.
Why "whole-body control" is a harder ML problem than it sounds
Coordinating legs and arms adds degrees of freedom, and it also changes the control problem: a legs-and-arms system has to reason about balance and momentum transfer in the same action space it uses for fine manipulation, meaning the same model has to be precise enough for five-finger dexterity and robust enough to not fall over when the arms' motion shifts the center of mass. That's a much wider dynamic range than either sub-problem alone, and it's the reason most prior systems punted on it by keeping the two separate. Multi-robot collaboration, which DeepMind also claims (ER 2 lets "different types of robots communicate and work together"), adds a further coordination layer on top of that, since now the action space includes reasoning about what another agent's body is doing, not just your own.
Atlas, AgiBot and OpenAI
The release lands in a crowded year. Boston Dynamics said in January that all Atlas deployments are fully committed for 2026, with fleets going to Hyundai and Google DeepMind; AgiBot rolled out its 10,000th humanoid in March, with the jump from 5,000 to 10,000 taking about three months after roughly three years to reach 5,000; and OpenAI's Sam Altman said on September 3 that the company "will definitely do a humanoid," though with no ship date or prototype announced. The industry's own framing has shifted from "humanoid robot" as the headline to "physical AI," with the emphasis on generalist models that adapt to new tasks and bodies from relatively little data (On-Device 2's fewer-than-200-example claim fits that trend) rather than robots pre-programmed move by move for one specific task.
What I'd look for in the next release
Walking and gripping a cup were both within reach of 2025-era systems. The test for the next release is whether one model (or a tightly coupled pair, as here) decides both at once, since that coordination is where earlier systems fell short. DeepMind says this is the first time its models control entire humanoids, and the release comes with a clearly described three-model split and a safety report. Both can be read and checked, and the safety report is where to start if you want the failure cases.