Skip to content
Gigantum.net
Streaming

Microsoft AI launches MAI-Transcribe-2-Streaming alongside multilingual MAI-Voice 2.1 models

Microsoft has launched MAI-Transcribe-2-Streaming model, with Microsoft AI CEO Mustafa Suleyman, calling it the most accurate real-time audio

· 438 words

Microsoft has launched MAI-Transcribe-2-Streaming model, with Microsoft AI CEO Mustafa Suleyman, calling it the most accurate real-time audio transcription model in the world. Unveiled alongside two companion speech generation models – MAI-Voice-2.1 and the high-speed MAI-Voice-2.1-Flash – the release is designed to provide developers with responsive, low-cost building blocks for building autonomous conversational voice agents.According to Suleyman, the new voice architecture delivers inference speeds that are 55% faster and operational costs that are 60% cheaper than comparable offerings like ElevenLabs. He said in a post on X (formerly Twitter): “We're launching the most accurate real time transcription model in the world... #1 !!! 55% faster and 60% cheaper than ElevenLabs. Come build agents on our platform!”Real-time speech recognition: MAI-Transcribe-2-StreamingThe headline addition to the company’s audio lineup, MAI-Transcribe-2-Streaming, provides continuous, low-latency live speech-to-text across 60 languages with built-in automatic language detection. Key performance benchmarks and features include top benchmark ranking as the model has secured the No. 1 spot for accuracy across both partial and completed transcripts on the independent benchmark platform Artificial Analysis.It sits on the Pareto frontier for accuracy versus latency, eliminating the historical tradeoff where high precision required sluggish processing times. Instead of pausing until a sentence ends, the model outputs initial text hypotheses within roughly 100 milliseconds of receiving incoming audio. It refines words as additional context arrives and finalizes stable text immediately.By surfacing partial transcripts as words are spoken, voice bots can trigger API tools and start logical reasoning before the user finishes speaking. For live captioning and automated dictation workflows, internal tests show text appears on-screen twice as fast as the nearest market competitor.Natural Multilingual Output: MAI-Voice-2.1To pair with speech recognition, the team introduced MAI-Voice-2.1, its flagship multilingual text-to-speech engine covering 23 languages and 26 regional locales.The system's standout capability is cross-lingual identity preservation. A single synthetic voice persona can switch seamlessly between diverse languages, such as English, Mandarin and German, without carrying a foreign accent over from one language to another. Instead, the synthesized speaker modifies actual local pronunciation and retains a unique vocal identity.That means brands can deploy uniform virtual ambassadors worldwide, educational sites can switch languages while keeping the same instructors, and customer support agents can respond in the caller’s language without missing a beat.MAI-Voice-2.1-Flash Low-Latency ScaleMAI-Voice-2.1-Flash was introduced by the company for mission-critical, high-capacity applications that demand immediate response times.The Flash version offers the same 23-language voice consistency as the base model, but it is optimized for high-throughput enterprise workloads, and can produce 45 seconds of synthesized audio with an end-to-end latency of just 150 milliseconds.You use AI every day. Now get your AI Quotient. Take the AIQ test.

Gathered from external sources. Rights to this text belong to whoever originally published it.

Friday, October 2, 2026

© 2026 Gigantum.net. Content gathered automatically from external sources; rights to each text belong to whoever originally published it.