Gemini 2.5 Flash Lite Preview 09.2025
Google (Gemini)
At a glance
- Price per 1M tokens (input)
- $0.10
- Price per 1M tokens (output)
- $0.40
- Context
- 1,048,576 tokens
- Free access
- yes (see below)
Refreshed daily; data verified August 16, 2026. Published August 12, 2026.
Gemini 2.5 Flash Lite Preview is aimed at developers and product teams that need quick visual analysis combined with logical reasoning without heavy infrastructure. It handles image interpretation, step‑by‑step problem solving, and dynamic function execution through tool calling, which lets the model invoke external code or APIs as part of a response. Google’s approach blends large‑scale multimodal training with a lightweight inference design, keeping latency low while preserving the depth of its research‑grade models. The result is a practical assistant for workflows that mix pictures, data queries, and procedural logic.
Specifications & pricing
| Input (per 1M tokens) | $0.10 |
|---|---|
| Output (per 1M tokens) | $0.40 |
| Cache read (per 1M tokens) | $0.01 |
| Context window | 1,048,576 tokens |
| Max output | 65,535 tokens |
| Capabilities | images, tool calling, reasoning |
LiteLLM community dataset (MIT), verified August 16, 2026. Official Google (Gemini) pricing.
Where to try Gemini 2.5 Flash Lite Preview 09.2025 for free
- Google (Gemini) offers a free chat — Gemini (free plan). A vendor's free chat may run a different model from the same family — the exact model is not guaranteed.
- 🎁 On our promo-codes page: Gemini for free: access and discounts.
Frequently asked questions
What kinds of business problems benefit most from this model?+
Tasks that require understanding of visual content together with structured reasoning, such as product catalog tagging, quality inspection, or generating reports that combine images and text, are a good fit.
How do I start using the model in my existing applications?+
You connect through the standard API endpoint, provide an image or text prompt, and optionally describe the functions you want the model to call; the service returns a response that may include a function call payload you can execute.
How does this version differ from the larger Gemini models?+
The Flash Lite variant is optimized for speed and lower compute, sacrificing some depth of knowledge and the ability to handle extremely large contexts, while still preserving multimodal perception and reasoning capabilities.
Compare with others

Org chart: how to move your company onto AI
A practical map: which company roles and processes AI agents can take over, where to start, and in what order to roll it out.