Back to AI models

Llama 3.2 90B Vision Instruct

Meta’s large open vision-language model for understanding images and documents.

Llama 3.2 Multimodal LLM Open weights Open weights: Yes API: Via Meta and partners

Plain-English Summary

Meta’s large open vision-language model for understanding images and documents.

Technical Notes

90B multimodal instruction model with image and text input, text output, and 128K context.

Best For

Document vision, image understanding, self-hosted multimodal RAG

Watch Out For

Large hardware requirement; primarily understanding rather than image generation.

Model Specs

Maker
Meta
Country
USA
Family
Llama 3.2
Type
Multimodal LLM
Release date
2024-09-25
Date confidence
Exact
Status
Open weights
Context window
131,072 tokens
Max output
Not listed
Parameters
90B
Active parameters
90B
Architecture
Vision-language transformer
Reasoning
General reasoning
Tool calling
Host dependent
Structured output
Host dependent
API available
Via Meta and partners
Open weights
Yes
Self-hostable
Yes
License
Llama 3.2 Community License
Inputs
Text, image
Outputs
Text
Approx. price
See provider pricing
How to access
Provider website
Model / API ID
Llama-3.2-90B-Vision-Instruct
Knowledge cutoff
Not publicly disclosed
Fine-tuning
Fine-tuning, quantization, self-hosting
Completeness
95 %
Verified
Jul 18, 2026
Snapshot date
Jul 18, 2026
Source
Official source