Plain-English Summary
A live conversational model that can listen, see, speak, and respond in real time.
A live conversational model that can listen, see, speak, and respond in real time.
A live conversational model that can listen, see, speak, and respond in real time.
Streaming native-multimodal model accepting text, images, audio, and video and producing text or audio with 131K context.
Voice agents, live assistants, video understanding, interactive experiences
Preview; real-time systems require careful latency, interruption, and safety design.