Google DeepMind has released EmbeddingGemma 2, a 740-million-parameter model that unifies text, images, audio, and video into a shared embedding space while remaining efficient enough for on-device inference. The model a…
#Multimodal Models
Models · Open in RealityHacker · RSS
DeepMind Releases EmbeddingGemma 2 for On-Device Multimodal Embeddings
Gemini 3.8 Live Adds Real-Time Animated Persona for Enterprise Conversations
Google DeepMind has introduced Gemini 3.8 Live with Live Avatar, a feature that adds a real-time animated visual persona to its conversational AI. The system synchronizes lip movements and facial expressions with speech,…