Plain-English Summary
Meta’s large open vision-language model for understanding images and documents.
Meta’s large open vision-language model for understanding images and documents.
Meta’s large open vision-language model for understanding images and documents.
90B multimodal instruction model with image and text input, text output, and 128K context.
Document vision, image understanding, self-hosted multimodal RAG
Large hardware requirement; primarily understanding rather than image generation.