LLMgram · AI News · 2026-08-10

NVIDIA Releases NemotronLabs VoiceChat 11B Open Full-Duplex Speech-to-Speech Model

NVIDIA Releases NemotronLabs VoiceChat 11B Open Full-Duplex Speech-to-Speech Model

NVIDIA added NemotronLabs VoiceChat 11B to the open voice stack, advertising full-duplex speech-to-speech interaction with about 448 ms turn-taking and live tool calling according to MarkTechPost. That combination targets agents that can listen, speak, and invoke tools in one conversational loop rather than chaining separate speech modules. The timing sits within a wider NVIDIA speech push that also includes NeMo-Speech.cpp and GGUF quantized ASR, TTS, and codec models surfaced on Reddit, plus OpenAI's recent GPT-Live write-up on continuous full-duplex voice without a separate turn detector. Operators should treat the latency figure as vendor-reported from early coverage, not yet validated by independent benchmarks or widely shared production deployments.

Sources

NVIDIA Releases NemotronLabs VoiceChat 11B Open Full-Duplex Speech-to-Speech Model

NVIDIA Releases NemotronLabs VoiceChat 11B Open Full-Duplex Speech-to-Speech Model

NVIDIA releases NemotronLabs VoiceChat 11B, an open full-duplex speech-to-speech model with 448 ms latency and live tool calling. MarkTechPost describes an open full-duplex speech-to-speech model with roughly 450 ms turn-taking.

Key takeaway

Open full-duplex voice is now a measurable builder surface: NVIDIA's 11B release pairs sub-500 ms turn-taking with live tool calling in an open weights package.

What happened

MarkTechPost reports that NVIDIA released NemotronLabs VoiceChat 11B, an open full-duplex speech-to-speech model with about 448 ms turn-taking latency and live tool calling.

The coverage frames the model as a high-performance open alternative for responsive voice agents, amid parallel NVIDIA speech moves including NeMo-Speech.cpp with GGUF-quantized ASR, TTS, and codec models discussed on Reddit and OpenAI's GPT-Live full-duplex voice architecture write-up.

Evidence

  • NVIDIA released NemotronLabs VoiceChat 11B as an open full-duplex speech-to-speech model with about 448 ms latency and live tool calling.

    MarkTechPost · attributed

    NVIDIA releases NemotronLabs VoiceChat 11B, an open full-duplex speech-to-speech model with 448 ms latency and live tool calling.

  • MarkTechPost describes roughly 450 ms turn-taking for the VoiceChat 11B release.

    MarkTechPost · attributed

    MarkTechPost describes an open full-duplex speech-to-speech model with roughly 450 ms turn-taking.

  • NVIDIA released NeMo-Speech.cpp with GGUF-quantized speech models for on-device ASR, TTS, and codec inference.

    r/LocalLLaMA Top · attributed

    NVIDIA has released NeMo-Speech.cpp, a C++ inference engine for its speech models, alongside GGUF-quantized versions of Magpie-TTS, Nemotron ASR, and Parakeet models.

  • OpenAI built GPT-Live as a full-duplex realtime voice system that removes the turn detector from the audio path.

    LLMgram AI News · attributed

    The system removes the turn detector from the audio path so the model can listen and speak at once, and can call frontier models such as GPT-5.5 for deeper reasoni

Why it matters

Teams building voice agents can prototype responsive, tool-enabled conversations without committing to proprietary realtime APIs, if the reported latency holds in their deployment stack.

Limits and uncertainties

MarkTechPost coverage is early reporting and does not cite independent latency benchmarks for VoiceChat 11B.

The packet does not provide production deployment results or hardware-specific latency validation beyond the announced figures.

Practical implications

Evaluate NemotronLabs VoiceChat 11B for open full-duplex voice agent prototypes that need live tool calling.

Compare the open NVIDIA speech path, including NeMo-Speech.cpp and GGUF quantized models, against closed full-duplex stacks such as OpenAI GPT-Live.

What to watch

Independent benchmark results confirming the reported 448 ms turn-taking latency on representative hardware.

Public model weights, integration guides, and live tool-calling examples tied to VoiceChat 11B.

Sources

LLMgram editorial selection and synthesis · @llmgram. LLMgram is not the original publisher of this information.
Continue on LLMgram: Open in AI Signal →
Original reporting: NVIDIA Releases NemotronLabs VoiceChat 11B: An Open Full-Duplex Speech-to-Speech Model with ~450 ms Turn-Taking and Live Tool Calling