Grok Vision Beta
xAI
At a glance
- Price per 1M tokens (input)
- $5.00
- Price per 1M tokens (output)
- $15.00
- Context
- 8,192 tokens
- Free access
- yes (see below)
Refreshed daily; data verified August 15, 2026. Published August 12, 2026.
Grok Vision Beta is built for teams that need a multimodal model capable of interpreting images and acting on that understanding through tool calling, which lets the model invoke external functions or APIs to complete tasks. It suits use cases like visual document analysis, UI screenshot inspection, and automated workflows where an image triggers a downstream action. The vendor’s approach emphasizes tight integration with real-world systems, prioritizing practical utility over open-ended conversation. This makes it a fit for developers and operations teams who want to embed vision-driven automation into their existing infrastructure.
Specifications & pricing
| Input (per 1M tokens) | $5.00 |
|---|---|
| Output (per 1M tokens) | $15.00 |
| Context window | 8,192 tokens |
| Max output | 8,192 tokens |
| Capabilities | images, tool calling |
LiteLLM community dataset (MIT), verified August 15, 2026. Official xAI pricing.
What Grok Vision Beta would cost on your workload — run it through the cost calculator →
Where to try Grok Vision Beta for free
- xAI offers a free chat — Grok (limited free access). A vendor's free chat may run a different model from the same family — the exact model is not guaranteed.
Frequently asked questions
What is this model actually good for in a business context?+
It excels at tasks that combine visual input with structured output or actions. Examples include extracting data from charts or receipts, validating interface layouts against design specs, and routing support tickets based on attached screenshots. The tool calling capability means it can hand off results to your internal systems, like a database or ticketing platform, without a human in the loop.
How do I get started using Grok Vision Beta?+
You access it through the xAI API, same as other Grok models. You send an image along with your prompt, and if you want it to trigger a function, you define the available tools in your request. The model will respond with a structured call to the appropriate tool, which your application then executes. Check the API documentation for request formatting and supported image types.
How does Grok Vision Beta differ from the other Grok models in the lineup?+
The core difference is the vision input: this model accepts images as part of the prompt, whereas the text-only Grok models do not. It also has a stronger emphasis on tool calling, meaning it is optimized to produce machine-readable function calls rather than purely conversational replies. If your workflow is text-only or purely chat-based, a sibling model may be more appropriate.
What are the main limitations I should plan around?+
As a beta, it may show inconsistent performance on highly complex or ambiguous images, such as dense scientific figures or heavily occluded scenes. It also requires you to design your tools and prompts carefully; the model will not infer your API schema on its own. Finally, it is not a general-purpose vision-language assistant for open-ended Q&A, so you should scope it to defined tasks rather than broad analysis.
Compare with others

Org chart: how to move your company onto AI
A practical map: which company roles and processes AI agents can take over, where to start, and in what order to roll it out.