Build Low-Latency Multilingual Voice Agents: Open Weights & Full Deployment Control with NVIDIA Magpie TTS
A Blog post by NVIDIA on Hugging Face
Modèles combinant texte, image, audio et vidéo.
Thème · 33 articles
A Blog post by NVIDIA on Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Shieldstral introduces a 3B open-weights multimodal safety classifier that outperforms models up to 7x its size.
Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications.
From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasks.
We’re introducing new Gemini models, including Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber.
The most powerful AI platform for enterprises. Customize, fine-tune, and deploy AI assistants, autonomous agents, and multimodal AI with open models.
Scale your ideas with Nano Banana 2 Lite, our fastest, most cost-efficient Gemini Image model, and Gemini Omni Flash for high-quality video and conversational editing.
The most powerful AI platform for enterprises. Customize, fine-tune, and deploy AI assistants, autonomous agents, and multimodal AI with open models.
An overview of Gemma 4 12B, a model designed to bring high-performance multimodal intelligence directly to your laptop.
Innovations for global enterprises solving the world’s hardest problems.
A new class of AI models that predict the behavior of physical systems, powering the engineers and hardware products of tomorrow.