Key Takeaways

- Gemini Robotics 2 is a vision-language-action model that can control robots of any form factor, from tabletop arms to full-body humanoids
- DeepMind also released Gemini Robotics ER 2, an embodied reasoning model available now in Google AI Studio
- Developers can apply for early access to Gemini Robotics 2 through a waitlist
Google DeepMind has released Gemini Robotics 2, a vision-language-action (VLA) model designed to serve as a general-purpose intelligence layer for robots. The model can control systems ranging from small tabletop arms to full-body humanoids, and developers can now apply for early access through a waitlist.
Alongside it, DeepMind shipped Gemini Robotics ER 2, an embodied reasoning model that handles higher-level decision-making. ER 2 is already available in Google AI Studio, replacing the ER 1.6 version released in April.
What is a vision-language-action model?
VLA models combine three capabilities that robots need to operate in the physical world. They process visual input (image recognition), understand instructions (language processing), and generate movement commands (action control). The result is a single model that can see, reason about what it sees, and act on that reasoning.
This architecture differs from traditional robotics pipelines, which typically chain together separate systems for perception, planning, and control. Each handoff between systems introduces latency and potential errors. A VLA model collapses these stages into one inference pass.
DeepMind describes Gemini Robotics 2 as capable of full-body movement control, fine motor tasks like manipulation, and multi-robot coordination. The claim that one model can handle this range of form factors is significant. Most existing robotics models are trained for specific hardware configurations and struggle to generalize across platforms.
Embodied reasoning: the ER 2 layer
Gemini Robotics ER 2 sits above the action model in the control stack. Where the VLA model handles moment-to-moment movement, the ER model handles higher-level reasoning: understanding the physical context, deciding what actions make sense, and sequencing multi-step tasks.
Think of it as the difference between knowing how to move your arm and deciding whether to pick up a cup or a pen. The ER model handles the latter. DeepMind uses the term "embodied reasoning" to describe this capability of understanding and acting within physical space.
ER 2 is available now in Google AI Studio, which lowers the barrier for developers who want to experiment. The previous version, ER 1.6, shipped just three months ago. That update cadence suggests DeepMind is iterating fast on this layer.
Who gets access and when
The two models have different access paths. Gemini Robotics ER 2 is available immediately through Google AI Studio, which most developers with Google Cloud accounts can access today. No waitlist, no application process.
Gemini Robotics 2, the full VLA model, requires applying through a waitlist. DeepMind has not published criteria for selection, but early access programs like this typically prioritize partners with existing hardware platforms and clear deployment use cases.
DeepMind has not disclosed pricing. Given that Google AI Studio already hosts other Gemini models with usage-based pricing, a similar structure is likely for ER 2. The VLA model, which generates continuous action outputs, could require different pricing mechanics tied to inference volume or control loop frequency.
What DeepMind is not saying
The announcement is light on technical details. DeepMind has not published model architecture, parameter counts, training data, or benchmark results. For a research lab that typically ships detailed papers alongside product releases, the absence is notable.
We also do not know the inference latency. Robotics control loops typically run at 100Hz or higher for smooth movement. If Gemini Robotics 2 cannot hit those frequencies, it may be limited to slower, coarser tasks. The claim that it handles "fine motor tasks" suggests acceptable latency, but no numbers are provided.
Hardware requirements are another gap. Running a large multimodal model on a robot means either onboard compute (expensive, power-hungry) or a cloud connection (latency, reliability risk). DeepMind does not specify which architecture they assume.
Finally, the safety story is thin. Robots acting in physical environments can cause real harm. DeepMind mentions coordination and control but says nothing about safety constraints, operational limits, or how the model handles edge cases.
The competitive landscape
DeepMind is not alone in pursuing foundation models for robotics. OpenAI has invested heavily in robotics through partnerships, though it has not shipped a comparable VLA model. Meta's FAIR lab has published research on embodied AI but lacks a product-grade offering.
Startups are moving faster on narrow verticals. Companies like Covariant (warehouse automation), Figure (humanoid robots), and 1X Technologies (general-purpose humanoids) are deploying systems in production. These players often use proprietary models tuned for specific hardware.
DeepMind's pitch is generality: one model that works across form factors. If the claim holds, it could accelerate robotics development by letting teams focus on hardware and applications rather than training custom perception and control systems. The trade-off is dependence on Google's infrastructure and whatever pricing and access terms emerge.
Practical implications for builders
For teams building robotics applications today, the ER 2 release is the immediate opportunity. You can start testing embodied reasoning capabilities against your use cases in Google AI Studio without hardware integration. This lets you evaluate whether the model's spatial understanding and task planning fit your requirements before committing to the full VLA stack.
The VLA model is harder to evaluate without access. If your roadmap includes new hardware platforms or you are currently maintaining separate perception and control systems, joining the waitlist makes sense. But plan for uncertainty: DeepMind has not committed to timelines, pricing, or even the supported hardware interfaces.
The release also changes the build-versus-buy calculus. Training custom VLA models requires massive compute budgets and robotics data that most teams do not have. If Gemini Robotics 2 performs as claimed, the cost equation shifts toward using a foundation model and focusing engineering effort on application-layer work.
Practical framework for evaluating and deploying AI models in production systems
Logicity's Take
DeepMind is positioning Gemini Robotics 2 as the Android of robotics: a general-purpose intelligence layer that any hardware maker can adopt. The strategy echoes Google's mobile playbook, and the business logic is similar. Control the AI layer, and you control the platform. For robotics teams, this creates a classic platform dependency question. The model could accelerate development significantly, but you are building on Google's terms, and those terms are not yet defined. The smart move is to test ER 2 now while it is open, prototype against the VLA model if you get access, but maintain fallback options until pricing and reliability are proven in production.
Frequently Asked Questions
How do I get access to Gemini Robotics 2?
Apply through the waitlist on DeepMind's site. Selection criteria are not public, but expect priority for teams with existing hardware platforms.
Can I use Gemini Robotics ER 2 today?
Yes. ER 2 is available now in Google AI Studio with no waitlist required.
What hardware does Gemini Robotics 2 support?
DeepMind claims support from tabletop arms to humanoids, but has not published specific hardware requirements or supported platforms.
Need Help Implementing This?
If you're evaluating Gemini Robotics for your hardware platform or exploring embodied AI integration, Logicity's team can help you scope requirements and build a proof of concept. Reach out to discuss your use case.
Source: The Decoder / Matthias Bastian
Huma Shazia
Senior AI & Tech Writer
Produced with AI assistance and reviewed by the Logicity editorial team. Learn more in our Editorial Policy.
Related Articles
More in AI & Machine Learning
Bezos AI Lab Gets $10B: What Project Prometheus Means
Jeff Bezos is closing a $10 billion funding round for Project Prometheus, an AI lab focused on physics-based AI for manufacturing and engineering. With a $38 billion valuation and backing from JPMorgan and BlackRock, this signals a major shift in enterprise AI investment toward industrial applications.

Kimi K2.6 Open-Weight AI: 300 Agents at a Fraction of the Cost
Moonshot AI's Kimi K2.6 matches GPT-5.4 and Claude Opus 4.6 on coding benchmarks while running 300 parallel agents. For businesses locked into expensive API contracts, this open-weight model could slash AI infrastructure costs while delivering enterprise-grade automation.




