Gemini Robotics 2 from Google DeepMind manages entire humanoid robots.
Google DeepMind aims to develop a single AI brain that can operate all robots, and it has now taught this brain to utilize its entire body. The company has introduced Gemini Robotics 2, a series of models capable of controlling a humanoid robot from its feet to its fingertips. This system can also manage multiple machines simultaneously and adapt to a new robot body within a few hours.
This development represents a significant advancement in what the industry refers to as “physical AI,” with Wired describing it as a genuine step toward “physical AGI.” The concept is straightforward: the same type of model that composes your emails is now learning to navigate cluttered environments and tidy them up.
The key advancement is control. Previous robot models from DeepMind mainly focused on operating the upper body for tasks on tables. However, Gemini Robotics 2 manages the entire robot. It can make a humanoid walk, crouch, stretch, and maintain balance while handling objects in confined spaces designed for humans.
In one demonstration, Apptronik’s Apollo 2 robot received a simple command: place the watering can in the green bin on the lowest shelf. It walked to a table, picked up the can, moved to the shelves, crouched down, and placed it inside. Although this may seem trivial, coordinating the legs, torso, arms, and hands from a single instruction is quite complex.
Dexterity has also seen significant improvement. The model can operate Apollo's five-fingered, 22-joint hand to tie knots and seal ziplock bags, and it can also manage simpler two-fingered grippers on other systems. “Our aim is to integrate AI into the physical world and to develop an intelligence layer that can be utilized by every robot,” stated Carolina Parada, DeepMind’s robotics lead.
The release encompasses three distinct models. Gemini Robotics 2 is a vision-language-action model that translates what a robot perceives into motor commands, enabling it to carry out physical tasks.
Gemini Robotics ER 2 serves as the reasoning layer, acting as a high-level brain that orchestrates multi-step tasks. In its developer presentation, Google demonstrated its ability to track its progress via a live video stream. It can utilize tools like Google Search and even command a Boston Dynamics Spot robot to retrieve a snack. It also enables collaboration between different robots, such as a wheeled machine and a humanoid sharing responsibilities. Developers can currently access it through the Gemini API and Google AI Studio.
The third model, On-Device 2, operates independently without the need for an internet connection. It can be adapted to a completely new robot body with under 200 examples and just a few hours of training. This capability is significant, as transferring a learned skill from one machine to another has long been a considerable challenge in robotics.
Despite these advancements, DeepMind has been candid about its limitations. Internal assessments reveal that the system can unscrew a light bulb 92% of the time, but it struggles with more intricate tasks. For example, the rates for tying a trash bag stand at 44%, while sealing a ziplock bag is at 40%.
The robots also tend to move slowly, taking time to deliberate over actions that humans execute instinctively. Kanishka Rao, a director at DeepMind, emphasized that achieving true dexterity is still a distant objective, noting that robots currently learn much less efficiently than humans, who typically adjust after one or two errors.
This issue is a common challenge in the field, as competing efforts from robotics foundation-model startups and dexterity projects involving other humanoids consistently encounter the same obstacles. While the demonstrations are impressive, the machines are still a long way from being ready for domestic use.
With robots becoming capable of moving in proximity to people, DeepMind has prioritized safety. The Gemini Robotics ER 2 model is touted as the safest yet, as it is more adept at detecting nearby individuals and stopping until the area is clear before resuming its tasks.
Additionally, the company has introduced a benchmark named ASIMOV-Agentic, which evaluates whether the reasoning model will reject unsafe commands from the action model, flag impossible tasks, or request human assistance. Naming a robot safety benchmark after Isaac Asimov is fitting, given the ongoing concern about machines acting on incorrect instructions.
There is also a challenging backdrop to consider. Google develops the software but does not manufacture the robots, and geopolitical tensions are affecting hardware availability. As Axios reported, the US has initiated bans on future sales of robots made in China due to security concerns. Many of the robots that this software could potentially operate may be constructed in China.
Google is collaborating with Western partners like Apptronik, Boston Dynamics, and Agile Robots, alongside more than 100 trusted testers. Competing firms, including OpenAI and Nvidia, are also developing robot models with similar aspirations: a single model to operate across various hardware.
DeepMind is cautious to refer to this as a milestone rather than a conclusion. Gemini Robotics 2 enhances
Other articles
Gemini Robotics 2 from Google DeepMind manages entire humanoid robots.
Google DeepMind's Gemini Robotics 2 provides a single AI model with full-body control over humanoid robots, enabling them to collaborate as a team, although their dexterity remains behind.
