Gemini Robotics Er 2 Streaming Preview
At a glance
- Price per 1M tokens (input)
- $1.00
- Price per 1M tokens (output)
- $5.00
- Context
- 131,072 tokens
- Free access
- yes (see below)
Refreshed daily; data verified October 1, 2026. Published October 1, 2026.
Gemini Robotics Er 2 Streaming Preview is designed for developers and engineers building intelligent systems that require real-time perception and actionable reasoning. It is well-suited for tasks such as interpreting visual inputs from cameras or sensors, invoking external tools or APIs based on contextual understanding, and performing multi-step logical reasoning to guide robotic or automated workflows. Google’s approach emphasizes tight integration between vision, reasoning, and tool use within a unified streaming interface, enabling responsive, context-aware behavior without requiring complex orchestration layers.
Specifications & pricing
| Input (per 1M tokens) | $1.00 |
|---|---|
| Output (per 1M tokens) | $5.00 |
| Context window | 131,072 tokens |
| Max output | 65,536 tokens |
| Capabilities | images, tool calling, reasoning |
LiteLLM community dataset (MIT), verified October 1, 2026. Official Google (Gemini) pricing.
Where to try Gemini Robotics Er 2 Streaming Preview for free
- Google (Gemini) offers a free chat — Gemini (free plan). A vendor's free chat may run a different model from the same family — the exact model is not guaranteed.
- 🎁 On our promo-codes page: Gemini for free: access and discounts.
Frequently asked questions
What kinds of tasks is this model particularly effective for?+
This model excels at processing visual data to understand scenes or objects, then using that understanding to decide which external functions to call — such as moving a robotic arm, querying a database, or adjusting a system setting — followed by reasoning about the outcome to determine next steps.
How does a developer begin using this model in an application?+
Access is provided through Google’s Vertex AI platform, where the model can be invoked via API with inputs including image data and task prompts; developers define available tools as functions with clear schemas, and the model returns structured calls to execute them.
How does this model differ from other Gemini variants like Gemini Pro or Ultra?+
While Pro and Ultra focus on broad language and reasoning capabilities, this variant is optimized for low-latency vision-tool-reasoning loops in embodied or interactive environments, with streaming output designed for continuous interaction rather than batch processing.
What are the current limitations users should be aware of?+
The model may struggle with highly ambiguous or low-quality visual inputs, and tool calling depends on the clarity and correctness of the provided function definitions; it does not autonomously learn new tools or adapt to undefined actions without explicit programming.
Compare with others

Org chart: how to move your company onto AI
A practical map: which company roles and processes AI agents can take over, where to start, and in what order to roll it out.