Gemini Omni 1.1 Flash
Google (Gemini)
At a glance
- Price per 1M tokens (input)
- $1.50
- Price per 1M tokens (output)
- $9.00
- Context
- 131,072 tokens
- Free access
- yes (see below)
Refreshed daily; data verified September 1, 2026. Published September 1, 2026.
Gemini Omni Flash is designed for developers and product teams that need rapid visual analysis combined with logical reasoning. It excels at tasks such as interpreting documents that contain both text and images, generating step‑by‑step explanations, and supporting multimodal assistants. Google’s approach integrates its large‑scale language foundation with a dedicated vision pathway, allowing the model to process visual cues without separate preprocessing. The architecture emphasizes low latency, making it suitable for interactive applications where quick turn‑around is essential.
Specifications & pricing
| Input (per 1M tokens) | $1.50 |
|---|---|
| Output (per 1M tokens) | $9.00 |
| Context window | 131,072 tokens |
| Max output | 65,536 tokens |
| Capabilities | images, reasoning |
LiteLLM community dataset (MIT), verified September 1, 2026. Official Google (Gemini) pricing.
What Gemini Omni 1.1 Flash would cost on your workload — run it through the cost calculator →
Where to try Gemini Omni 1.1 Flash for free
- Google (Gemini) offers a free chat — Gemini (free plan). A vendor's free chat may run a different model from the same family — the exact model is not guaranteed.
- 🎁 On our promo-codes page: Gemini for free: access and discounts.
Frequently asked questions
What kinds of business problems benefit most from Gemini Omni Flash?+
The model is well suited for workflows that involve mixed media, such as processing scanned forms, analyzing product images alongside descriptions, and providing detailed reasoning in customer‑support chat that references visual content.
How do I start using the model in my application?+
You can access the model through Google’s cloud AI endpoint, authenticate with standard service credentials, and call the API with either text‑only or multimodal payloads. The SDK includes helper functions for handling image encoding and for streaming step‑by‑step responses.
How does Omni Flash differ from other Gemini models in the lineup?+
While the broader Gemini family focuses on pure language tasks, Omni Flash adds a high‑throughput vision encoder and prioritizes response speed, which makes it a better fit for real‑time visual question answering and interactive assistants.
What limitations should I consider before deploying this model?+
Current constraints include occasional confidence mis‑alignment on ambiguous images, support limited to common raster formats, and higher compute demand when processing large batches, which may affect cost and latency under heavy load.
Compare with others

Org chart: how to move your company onto AI
A practical map: which company roles and processes AI agents can take over, where to start, and in what order to roll it out.