Build Low-Latency Multilingual Voice Agents: Open Weights & Full Deployment Control with NVIDIA Magpie TTS
A Blog post by NVIDIA on Hugging Face
Reconnaissance et synthèse vocale, musique et audio.
Thème · 54 articles
A Blog post by NVIDIA on Hugging Face
A Blog post by Multiverse Computing on Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Computer-use AI agents struggle with multi-step workflows like email and customer support. Echoverse trains agents in realistic environments rather than simply providing more training tasks, helping them improve as the tasks, tests, and environments evolve…
Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications.
Cryptographic code supports vital protections in modern computing systems. Learn how a new method helps verify code as developers write it while preserving speed and adaptability as it gets implemented and evolves. The post Verifying Rust cryptography in…
A look at the built-in computer use tool in Gemini 3.5 Flash.
Gemini 3.5 Live Translate brings near real-time, natural speech translation to Google AI Studio, Google Translate and Google Meet.
An overview of Gemma 4 12B, a model designed to bring high-performance multimodal intelligence directly to your laptop.
Search Toolkit is a composable framework for building production search pipelines for AI applications.
We're releasing Stable Audio 3.0, a model family trained on fully licensed data, designed to be the foundation for what the audio community builds next.
Voxtral TTS: A frontier, open-weights text-to-speech model that’s fast, instantly adaptable, and produces lifelike speech for voice agents.