2026-07-30 · ← Radar
Gemini Robotics ER 2 shifts robots from executors to self-monitoring agents
Robots now watch what they actually did
Google DeepMind is launching Gemini Robotics ER 2, an upgraded model designed for physical agents. While previous systems mostly executed a specified sequence of movements blindly, the new version relies on continuous video processing. The model acts as a control layer that monitors the video feed from the environment in real time and checks whether the robotic arm actually reached its destination and if the task was successful.
Alongside video understanding, ER 2 brings better tool orchestration and the ability to coordinate multiple robots simultaneously in a shared space. Developers get access to the model via the Gemini API and Google AI Studio, albeit with an explicit ban on deployment in safety-critical areas like healthcare or transportation.
Upgraded navigation for industrial deployment
For integrators and industrial development teams, this changes the control architecture. Until now, they had to build complex state machines that handled every possible deviation. The model's ability to continuously monitor its own progress means control shifts from fixed scenarios to specifying a target state.
Shared spatial understanding also makes deploying multiple robots from different manufacturers cheaper and simpler. Instead of each machine operating in its own silo with a dedicated vision system, ER 2 can act as a central brain that delegates tasks across hardware based on visual input.
The ban highlights unreliability in critical moments
The restriction for healthcare and transportation clearly shows where the actual limits of today's robotics models lie. Video understanding works great in factory halls and warehouses, where a mistake means a dropped box, not a threat to human life.
Real-time video processing is also computationally extremely expensive. ER 2 must constantly evaluate huge amounts of visual data, which places heavy demands on local connectivity and Google's cloud infrastructure. In a real production environment where connection drops can occur, the model becomes a potential bottleneck.
Adoption outside Google's labs will decide
The proof of the model's viability will not be a demo video from DeepMind, but the willingness of hardware manufacturers to connect their machines to a third-party API. It will be crucial to watch whether large industrial companies agree to have their robots' logic driven by a Google cloud model, or if they prefer their own local, albeit less advanced, solutions.
Lilith's verdict
Google is not trying to sell a smarter camera. It wants to become the central operating system for every robotic arm in the factory, regardless of whose logo is stamped on the hardware.
I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.
Original source ↗ ↗