Microsoft AI just released MAI-Transcribe-2-Streaming, which is a brand-new model for real-time transcription. Microsoft says this model ranks first for accuracy on Artificial Analysis.
The system transcribes sixty languages. It delivers its first partial results in just over one hundred milliseconds. Microsoft says this speed lets voice agents respond while someone is still mid-sentence.
An hour of audio costs fifty-four cents at the introductory price. This promotional pricing runs through the end of the year.
Read nextCloudflare Releases Clef Decision Models to Challenge JevMicrosoft added two new text-to-speech models
Microsoft also launched two new text-to-speech models. MAI-Voice-2.1 speaks twenty-three languages in the same voice. It uses a native accent in each one.
The MAI-Voice-2.1-Flash variant hits a latency of one hundred fifty milliseconds. It costs fifteen dollars per million characters instead of twenty-two dollars.

Both voice models can clone a voice from just a few seconds of reference audio. Built-in safeguards are meant to prevent misuse.
The models are available through Microsoft Foundry and the MAI Playground. The two voice models are also available on OpenRouter, among other platforms.
In one test, about half of the four thousand participants thought the generated voices belonged to a real person.



