Human foundation model
AI foundation modelA real-time audiovisual conversational model that listens, reasons and speaks simultaneously, reading tone, facial expression, gesture and hesitation from live video and generating an expressive avatar response.
- Real-time emotion, expression and gesture detection from live video
- Autoregressive transformer over audiovisual tokens, extending LLM architectures to vision and audio
- Expressive avatar generation with dynamic facial expression, body language and tone
- Sub-500-millisecond response latency