Grok 4.20 Experimental Beta 0304
At a glance
- Price per 1M tokens (input)
- $1.25
- Price per 1M tokens (output)
- $2.50
- Context
- 1,000,000 tokens
- Free access
- yes (see below)
Refreshed daily; data verified September 25, 2026. Published September 25, 2026.
Grok 4.20 Experimental Beta 0304 by xAI is designed for developers and technical teams building applications that require visual comprehension and automated workflow integration. It excels at tasks such as interpreting diagrams, analyzing screenshots for debugging, and orchestrating multi-step processes through tool calling, which allows the model to invoke external functions or APIs to retrieve data or perform actions. Unlike more general-purpose models, this version emphasizes reasoning transparency by breaking down complex decisions into explicit steps, aiding in auditability and error tracing. The model reflects xAI’s focus on aligning advanced capabilities with practical utility in real-world software systems.
Specifications & pricing
| Input (per 1M tokens) | $1.25 |
|---|---|
| Output (per 1M tokens) | $2.50 |
| Cache read (per 1M tokens) | $0.20 |
| Context window | 1,000,000 tokens |
| Max output | 1,000,000 tokens |
| Capabilities | images, tool calling, reasoning |
LiteLLM community dataset (MIT), verified September 25, 2026. Official xAI pricing.
Where to try Grok 4.20 Experimental Beta 0304 for free
- xAI offers a free chat — Grok (limited free access). A vendor's free chat may run a different model from the same family — the exact model is not guaranteed.
Frequently asked questions
What types of tasks is this model particularly suited for?+
This model is well-suited for applications involving image understanding combined with reasoning, such as interpreting technical drawings, analyzing user interface elements for accessibility testing, or processing visual inputs in automated quality control pipelines. Its tool calling ability enables it to interact with databases, code repositories, or external services to gather information or trigger actions based on visual input. It is especially useful in workflows where decisions depend on both visual data and logical inference.
How does a developer begin using this model in an application?+
Developers can access the model through xAI’s API by specifying the model identifier in their requests, following the standard format for chat completions with support for image inputs and function definitions. The API accepts multimodal prompts where images are provided alongside text, and tool calls are structured as JSON schemas that the model can invoke. Documentation includes examples for setting up function calling and handling step-by-step reasoning outputs, enabling integration into existing development workflows.
How does this model differ from other versions in the Grok series?+
This experimental beta focuses on enhancing visual reasoning and tool interaction compared to earlier versions that may have prioritized text-only reasoning or broader knowledge recall. It places stronger emphasis on generating traceable, step-by-step logic chains when solving problems, which supports use cases requiring explainability. While sibling models might offer larger context or faster response times, this version is tuned for precision in multimodal task execution rather than scale.
What are the current limitations users should be aware of?+
As an experimental beta, the model may exhibit inconsistent performance on highly complex or ambiguous visual inputs, and tool calling reliability can vary depending on the complexity of the invoked functions. Step-by-step reasoning, while designed to improve transparency, does not guarantee correctness and should be validated in critical applications. Users are advised to test thoroughly in controlled environments before deploying in production systems due to the beta status.
Compare with others

Org chart: how to move your company onto AI
A practical map: which company roles and processes AI agents can take over, where to start, and in what order to roll it out.