Lyria 3.5 Clip Preview
At a glance
- Price per 1M tokens (input)
- free
- Price per 1M tokens (output)
- free
- Context
- 131,072 tokens
- Free access
- yes (see below)
Refreshed daily; data verified September 7, 2026. Published September 7, 2026.
Lyria 3.5 Clip Preview by Google (Gemini) is designed for developers and product teams building applications that require coherent, context-aware text generation from multimodal inputs. It supports tasks such as drafting responses, summarizing content, and generating structured outputs based on both textual and visual cues. What distinguishes this model is its integration within Google’s Gemini family, which emphasizes unified reasoning across modalities without requiring separate pipelines for different input types. This approach reduces complexity when handling inputs that combine text with images or other media.
Specifications & pricing
| Input (per 1M tokens) | free |
|---|---|
| Output (per 1M tokens) | free |
| Context window | 131,072 tokens |
| Max output | 8,192 tokens |
| Capabilities | text |
LiteLLM community dataset (MIT), verified September 7, 2026. Official Google (Gemini) pricing.
What Lyria 3.5 Clip Preview would cost on your workload — run it through the cost calculator →
Where to try Lyria 3.5 Clip Preview for free
- Google (Gemini) offers a free chat — Gemini (free plan). A vendor's free chat may run a different model from the same family — the exact model is not guaranteed.
- 🎁 On our promo-codes page: Gemini for free: access and discounts.
Frequently asked questions
What types of tasks is this model best suited for?+
It is effective for generating natural language responses when given text and image inputs together, such as describing visual content, answering questions about images, or creating captions that reflect both textual and visual context.
How does this model differ from other versions in the Gemini line?+
This preview version focuses on clip-based understanding, meaning it processes short sequences of visual and textual data in tandem, unlike text-only variants that do not incorporate image analysis.
What should users know about its limitations?+
As a preview release, it may not handle highly complex or ambiguous visual scenes with consistent accuracy, and performance can vary depending on the clarity and relevance of the provided image-text pairs.
How can a team begin using this model in their workflow?+
Access is typically provided through Google’s AI platforms with documented APIs; teams should start by testing with small, well-defined input pairs to evaluate output quality before scaling to production use.
Compare with others

Org chart: how to move your company onto AI
A practical map: which company roles and processes AI agents can take over, where to start, and in what order to roll it out.