Skip to content

Gemini Omni 1.1 Flash

Google (Gemini)

At a glance

Price per 1M tokens (input)
$1.50
Price per 1M tokens (output)
$9.00
Context
131,072 tokens
Free access
yes (see below)

Refreshed daily; data verified September 1, 2026. Published September 1, 2026.

Gemini Omni Flash is designed for developers and product teams that need rapid visual analysis combined with logical reasoning. It excels at tasks such as interpreting documents that contain both text and images, generating step‑by‑step explanations, and supporting multimodal assistants. Google’s approach integrates its large‑scale language foundation with a dedicated vision pathway, allowing the model to process visual cues without separate preprocessing. The architecture emphasizes low latency, making it suitable for interactive applications where quick turn‑around is essential.

Specifications & pricing

Input (per 1M tokens)$1.50
Output (per 1M tokens)$9.00
Context window131,072 tokens
Max output65,536 tokens
Capabilitiesimages, reasoning

LiteLLM community dataset (MIT), verified September 1, 2026. Official Google (Gemini) pricing.

What Gemini Omni 1.1 Flash would cost on your workload — run it through the cost calculator →

Where to try Gemini Omni 1.1 Flash for free

Frequently asked questions

What kinds of business problems benefit most from Gemini Omni Flash?+

The model is well suited for workflows that involve mixed media, such as processing scanned forms, analyzing product images alongside descriptions, and providing detailed reasoning in customer‑support chat that references visual content.

How do I start using the model in my application?+

You can access the model through Google’s cloud AI endpoint, authenticate with standard service credentials, and call the API with either text‑only or multimodal payloads. The SDK includes helper functions for handling image encoding and for streaming step‑by‑step responses.

How does Omni Flash differ from other Gemini models in the lineup?+

While the broader Gemini family focuses on pure language tasks, Omni Flash adds a high‑throughput vision encoder and prioritizes response speed, which makes it a better fit for real‑time visual question answering and interactive assistants.

What limitations should I consider before deploying this model?+

Current constraints include occasional confidence mis‑alignment on ambiguous images, support limited to common raster formats, and higher compute demand when processing large batches, which may affect cost and latency under heavy load.

Compare with others

Org chart: how to move your company onto AI

Org chart: how to move your company onto AI

A practical map: which company roles and processes AI agents can take over, where to start, and in what order to roll it out.