Plain-English Summary
A compact Phi model that can understand text, images, and audio in private or edge deployments.
A compact Phi model that can understand text, images, and audio in private or edge deployments.
A compact Phi model that can understand text, images, and audio in private or edge deployments.
Multimodal model accepting text, image, and audio with 131K input context and 4K output.
Document vision, speech and image understanding, private multimodal assistants
Primarily an understanding model; smaller capacity than large multimodal systems.