Back to AI models

Gemini 3.1 Flash Live Preview

A live conversational model that can listen, see, speak, and respond in real time.

Gemini Live Multimodal LLM Preview Open weights: No API: Yes

Plain-English Summary

A live conversational model that can listen, see, speak, and respond in real time.

Technical Notes

Streaming native-multimodal model accepting text, images, audio, and video and producing text or audio with 131K context.

Best For

Voice agents, live assistants, video understanding, interactive experiences

Watch Out For

Preview; real-time systems require careful latency, interruption, and safety design.

Model Specs

Maker
Google
Country
Not listed
Family
Gemini Live
Type
Multimodal LLM
Release date
2026-03-01
Date confidence
Month / latest update
Status
Preview
Context window
131,072 tokens
Max output
65,536
Parameters
Not disclosed
Active parameters
Not disclosed
Architecture
Proprietary streaming multimodal transformer
Reasoning
Yes
Tool calling
Yes
Structured output
Limited / endpoint dependent
API available
Yes
Open weights
No
Self-hostable
No
License
Proprietary
Inputs
Text, image, audio, video
Outputs
Text, audio
Approx. price
See provider pricing
How to access
Provider website or API
Model / API ID
gemini-3.1-flash-live-preview
Knowledge cutoff
2025-01
Fine-tuning
Prompting / RAG
Completeness
90 %
Verified
Jul 18, 2026
Snapshot date
Jul 18, 2026
Source
Official source