Back to AI models

Command A Vision

A Cohere model for understanding business documents, images, and mixed visual-text content.

Command Multimodal LLM GA Open weights: No API: Yes

Plain-English Summary

A Cohere model for understanding business documents, images, and mixed visual-text content.

Technical Notes

Text-and-image model with 128K context, 8K output, enterprise document understanding, and tool support.

Best For

Document processing, charts, scans, enterprise visual RAG

Watch Out For

Focused on understanding rather than image generation.

Model Specs

Maker
Cohere
Country
Canada
Family
Command
Type
Multimodal LLM
Release date
2025-07-01
Date confidence
Month / model ID
Status
GA
Context window
131,072 tokens
Max output
8,192
Parameters
Not disclosed
Active parameters
Not disclosed
Architecture
Proprietary vision-language transformer
Reasoning
General reasoning
Tool calling
Yes
Structured output
Yes
API available
Yes
Open weights
No
Self-hostable
Private deployment options
License
Proprietary
Inputs
Text, image
Outputs
Text
Approx. price
See provider pricing
How to access
Provider website or API
Model / API ID
command-a-vision-07-2025
Knowledge cutoff
Not publicly disclosed
Fine-tuning
RAG, tools, private deployment
Completeness
90 %
Verified
Jul 18, 2026
Snapshot date
Jul 18, 2026
Source
Official source