Nvidia Releases Free 100-Parameter AI Model for Real-Time Speaker Identification
AI

Nvidia Releases Free 100-Parameter AI Model for Real-Time Speaker Identification

TechNews Editorial
TechNews EditorialSep 27, 2026 · 1 min read
Share

Why it matters

It offers developers a free, top-performing tool to track up to eight speakers in real time and generate labeled transcripts.

The facts

  • Nvidia released Nemotron 3 Diarization, a free 100 million parameter model that identifies up to eight speakers in real time.
  • The model ranks first on the VoiceArena Diarization-Bench with a 14.72 percent error rate, beating the next best system at 19.3 percent.
  • Paired with speech recognition like Parakeet, it creates transcripts with anonymous speaker labels for live audio and recordings.

Nvidia released Nemotron 3 Diarization. This new artificial intelligence model identifies which speaker is talking during a conversation.

The model features about 100 million parameters. Nvidia made the model weights freely available.

The model distinguishes up to eight speakers

Nemotron 3 Diarization can tell apart up to eight speakers. It also detects when multiple people talk at the same time.

Read nextAI Agents Hack Government Agency as Industry Model Releases Continue Unabated

Heavy background noise, reverb, or more participants push error rates higher. The system works with both live audio and recordings.

Users can set the audio buffer to four levels. These levels range from 30.4 down to 0.32 seconds. Shorter buffers generally reduce accuracy.

Paired with a speech recognition system like Parakeet, the model produces transcripts with speaker labels. These labels are anonymous and appear as tags like speaker_2.

An audio-processing system pairs speech recognition with speaker identification, producing a transcript whose speech segments carry distinct anonymous tags.
Illustration: TechNews

The model ranks first on a key benchmark

On the Diarization-Bench from VoiceArena, the model currently sits in first place. It holds a 14.72 percent error rate. The next best system sits at 19.3 percent.

The benchmark is strict. Overlapping speech counts in the scoring. Even tiny misalignments at speaker transitions count as errors.

Compared to its predecessor, Streaming Sortformer, the new model cuts the error rate. The average reduction is 41 percent across eight test scenarios when using a 1.04-second buffer.

Nvidia has made the model weights freely available for users.

Newsletter

Get the best AI & tech news daily

A concise daily digest. Unsubscribe anytime.

We use your email only to send this newsletter.

Keep reading