Grok 4.20 Beta 0309
At a glance
- Price per 1M tokens (input)
- $1.25
- Price per 1M tokens (output)
- $2.50
- Context
- 1,000,000 tokens
- Free access
- yes (see below)
Refreshed daily; data verified September 25, 2026. Published September 25, 2026.
Grok 4.20 Beta 0309 by xAI is designed for developers and technical teams building applications that require visual understanding combined with actionable reasoning. It suits tasks such as interpreting diagrams or screenshots to trigger automated workflows, analyzing visual data for decision support, and executing multi-step logic where image input informs subsequent tool use. The model distinguishes itself through tight integration of image understanding with function calling, enabling a seamless flow from perception to action without requiring separate pipelines for vision and reasoning.
Specifications & pricing
| Input (per 1M tokens) | $1.25 |
|---|---|
| Output (per 1M tokens) | $2.50 |
| Cache read (per 1M tokens) | $0.20 |
| Context window | 1,000,000 tokens |
| Max output | 1,000,000 tokens |
| Capabilities | images, tool calling, reasoning |
LiteLLM community dataset (MIT), verified September 25, 2026. Official xAI pricing.
What Grok 4.20 Beta 0309 would cost on your workload — run it through the cost calculator →
Where to try Grok 4.20 Beta 0309 for free
- xAI offers a free chat — Grok (limited free access). A vendor's free chat may run a different model from the same family — the exact model is not guaranteed.
Frequently asked questions
What is tool calling and how does it work with this model?+
Tool calling allows the model to invoke external functions or APIs based on its reasoning. When given a task, it can determine which tool to use, prepare the necessary input, and execute the action — all while maintaining context from prior steps, including information derived from images.
How do I begin using this model in my application?+
Access is provided through xAI's API platform. Users send prompts that may include image data along with text, and the model responds with reasoning steps or tool invocation requests. Detailed integration guides and example code are available in the developer documentation.
How does this model differ from earlier versions in the Grok series?+
This version enhances the model's ability to process and reason over visual inputs while maintaining strong tool use performance. Earlier versions focused more on text-only reasoning or had less coordinated vision and action capabilities.
What are the current limitations of this model?+
The model may struggle with highly abstract or low-contrast images, and its reasoning depth can vary depending on the complexity of the task and the clarity of the visual input. It also depends on the availability and correct definition of external tools for successful execution.
Compare with others

Org chart: how to move your company onto AI
A practical map: which company roles and processes AI agents can take over, where to start, and in what order to roll it out.