Back to AI models

Phi-4-multimodal-instruct

A compact Phi model that can understand text, images, and audio in private or edge deployments.

Phi-4 Multimodal LLM GA / open weights Open weights: Yes API: Yes and self-hosted

Plain-English Summary

A compact Phi model that can understand text, images, and audio in private or edge deployments.

Technical Notes

Multimodal model accepting text, image, and audio with 131K input context and 4K output.

Best For

Document vision, speech and image understanding, private multimodal assistants

Watch Out For

Primarily an understanding model; smaller capacity than large multimodal systems.

Model Specs

Maker
Microsoft
Country
USA
Family
Phi-4
Type
Multimodal LLM
Release date
2025-02-26
Date confidence
Exact
Status
GA / open weights
Context window
131,072 tokens
Max output
4,096
Parameters
Not disclosed
Active parameters
Not disclosed
Architecture
Multimodal transformer
Reasoning
General reasoning
Tool calling
Host dependent
Structured output
Host dependent
API available
Yes and self-hosted
Open weights
Yes
Self-hostable
Yes
License
MIT
Inputs
Text, image, audio
Outputs
Text
Approx. price
See provider pricing
How to access
Provider website
Model / API ID
Phi-4-multimodal-instruct
Knowledge cutoff
Not publicly disclosed
Fine-tuning
Fine-tuning, quantization, self-hosting
Completeness
90 %
Verified
Jul 18, 2026
Snapshot date
Jul 18, 2026
Source
Official source