Ministral 3b 2512
Mistral AI
At a glance
- Price per 1M tokens (input)
- $0.10
- Price per 1M tokens (output)
- $0.10
- Context
- 131,072 tokens
- Free access
- yes (see below)
Refreshed daily; data verified August 30, 2026. Published August 30, 2026.
Ministral 3b 2512 is aimed at enterprises that need to combine visual analysis with automated workflow integration. It handles tasks such as image classification, visual question answering, and invoking external functions to retrieve or process data. Mistral AI builds the model with a focus on modular tool calling, a capability that lets the system decide when to run a function rather than relying on static prompts. The architecture balances language understanding with visual perception to support mixed‑modal applications without requiring separate models.
Specifications & pricing
| Input (per 1M tokens) | $0.10 |
|---|---|
| Output (per 1M tokens) | $0.10 |
| Context window | 131,072 tokens |
| Max output | 131,072 tokens |
| Capabilities | images, tool calling |
LiteLLM community dataset (MIT), verified August 30, 2026. Official Mistral AI pricing.
What Ministral 3b 2512 would cost on your workload — run it through the cost calculator →
Where to try Ministral 3b 2512 for free
- Mistral AI offers a free chat — Le Chat (free plan). A vendor's free chat may run a different model from the same family — the exact model is not guaranteed.
Frequently asked questions
What types of problems is this model best suited for?+
It excels at scenarios where image input must be interpreted and then used to trigger downstream actions, such as automated document processing, visual inspection, or customer support that includes photos.
How do I start using the model in my existing pipeline?+
You can access it through Mistral AI’s API, send a request that includes an image and optional function specifications, and handle the returned function call payload to integrate with your services.
How does Ministral 3b 2512 differ from other models in the Ministral family?+
Unlike the text‑only variants, this model incorporates a visual encoder and a dedicated tool‑calling module, so it can both see and act, whereas its siblings focus solely on language generation.
What are the main limitations I should be aware of?+
The model may struggle with highly specialized visual domains not represented in its training data, and tool calling depends on correctly defined function schemas; misaligned schemas can cause failures.
Compare with others

Org chart: how to move your company onto AI
A practical map: which company roles and processes AI agents can take over, where to start, and in what order to roll it out.