用全球最好的语音和流式转录模型构建你的智能体……现在可以在 @vercel 上轻松使用了 质量与速度均排名第一……比任何其他超大规模云厂商都便宜……比 Eleven Labs 便宜高达 60%
The @MicrosoftAI team is training excellent models. Excited to bring them to @vercel, day zero.
@mustafasuleyman · X
用全球最好的语音和流式转录模型构建你的智能体……现在可以在 @vercel 上轻松使用了 质量与速度均排名第一……比任何其他超大规模云厂商都便宜……比 Eleven Labs 便宜高达 60%
The @MicrosoftAI team is training excellent models. Excited to bring them to @vercel, day zero.
Microsoft AI has released MAI-Transcribe-2-Streaming, taking the #1 spot for Final Transcript accuracy and First Partial Transcript accuracy on AA-WER Streaming with 2.5% WER at 0.13s after end of speech MAI-Transcribe-2-Streaming is @MicrosoftAI's new streaming Speech to Text model, joining the non-streaming MAI-Transcribe-2. It leads streaming Final Transcript WER at 2.5%, ahead of the previous #1, SpaceXAI's Grok Voice Transcribe 2.0 at 2.7%, and returns that transcript in 0.13s rather than 0.49s. It is available for streaming transcription at $0.54 per hour of audio, at the higher end of pricing among the leading streaming models. Key takeaways ➤ Final Transcript: MAI-Transcribe-2-Streaming achieves 2.5% WER at 0.13s after end of speech, ranking #1 of 38 models. It is more accurate and faster than Grok Voice Transcribe 2.0 at 2.7% and 0.49s, Muse Voice Transcribe at 3.1% and 0.16s, and ElevenLabs Scribe v2 Realtime at 3.6% and 0.14s. It is also more accurate, though slightly slower, than Cartesia Ink Preview (external endpoints) at 3.1% and 0.11s ➤ First Partial Transcript: The model achieves 2.5% WER at 0.12s, ahead of Grok Voice Transcribe 2.0 at 3.4% and 0.49s, and ahead of Muse Voice Transcribe and ElevenLabs Scribe v2 Realtime on accuracy, both at 3.6%, and slightly faster than both at 0.12s versus 0.13s. It is more accurate but slower than Cartesia Ink-2 (external endpoints) at 4.0% and 0.07s ➤ Price: MAI-Transcribe-2-Streaming costs $0.54 per hour for streaming, or $9.00 per 1,000 minutes. This puts it level with Gemini 3.5 Transcribe Live at $9, above the $6.50 charged for ElevenLabs Scribe v2 Realtime and Deepgram Flux, more than twice Cartesia Ink-2 at $4, and three times Muse Voice Transcribe at $3. Non-streaming transcription costs $0.10 per hour, or $1.67 per 1,000 minutes Congrats to the @MicrosoftAI team on the launch! See more details below
Copilot 团队做得太棒了。Autopilot 非常酷。去看看!
Home. Code. Autopilot. https://x.com/i/article/2103304083172167680
这份跨党派的人类主义 AI 宣言中有很多非常好的提议。仍有一些值得我们讨论,但总体上是正确方向。我鼓励大家都去看一看。
I'm delighted to share that @mustafasuleyman, CEO of Microsoft AI, co-founder of Google DeepMind and Inflection AI, has signed the Pro-Human AI Declaration. If you too support it, please join him and over a million others by signing it here – the momentum is building! Let's build tools not beings & keep humans in charge. https://humanstatement.org
Generating high-quality images is cheaper and faster than ever. Muse Image, MAI-Image-2.6 and GPT Images 2.5 have substantially shifted the Text to Image Pareto frontiers for both price and speed in recent weeks.
这是一个非常直白且符合常识的观点:技术的目的是服务人类,加速人类繁荣。 任何无法实现这一目标的技术都是失败的,应当被拒绝。 我们还没有到那一步。但开始为这种可能性做准备是正确的。
Any pursuit of superintelligence has to be grounded in the core principle that if the AI we build is not helping humanity and under human control, it's not worth pursuing. We also need to accelerate and spread the benefits of AI, such that they are diffused broadly across countries, communities, and companies. This requires a frontier ecosystem in which both closed and open-source models can thrive. And for firms, it’s imperative that they retain full control over their unique and tacit knowledge. Every organization should be able to build its own continuous learning loop/hill climbing machine, without becoming dependent on any one model provider, and have the ability to embed its own knowledge into models and weights they control. So, in this context, we welcome the research, focus, and deliberate pacing needed to get alignment right as the design goal. We also welcome ideas like "embedded evaluators" and the broader efforts to develop the mechanisms to make this more than just talk. The key is that this cannot be controlled by a handful of entities, but must have broad representation across the ecosystem, countries, and fields, including academia. This is the approach we are taking: broad access and choice at every layer of the AI stack; enterprise control of learning loops and models; and the “Code of Conduct” that underlies our own first party MAI models that we’ll publish tomorrow for public consultation.
Microsoft's MAI-Image-2.6-Preview lands at #1 on the Artificial Analysis Image Editing Leaderboard and takes #2 in Text to Image MAI-Image-2.6 is the newest model in Microsoft AI's MAI-Image family, announced August 10. Microsoft highlights stronger text rendering, better portraits and 3D imagery, and more polished commercial and photorealistic outputs. Like the MAI-Image-2.5 family, it handles both text to image generation and image editing. In the Artificial Analysis Image Arena, MAI-Image-2.6-Preview debuts at #1 on the Image Editing Leaderboard, ahead of Microsoft's own MAI-Image-2.5-Pro, Reve 2.1, and OpenAI's GPT Image 2, and giving Microsoft the top two spots on the board. In Text to Image it takes #2, behind only OpenAI's GPT Image 2 and ahead of Reve 2.1. On our refreshed Text to Image taxonomy, MAI-Image-2.6-Preview takes the top spot on 5 of the 19 category leaderboards (Material, Knowledge, Frontier, Retail & Ecommerce, and Marketing & Advertising), ahead of GPT Image 2, which leads every other category. MAI-Image-2.6 extends a rapid run of strong image releases from Microsoft AI. MAI-Image-2.5 debuted at #2 in Text to Image in June. MAI-Image-2.5-Pro, launched July 23, took #1 in Image Editing when we published our results last week. MAI-Image-2.6 now takes that top spot from its own sibling, and sits at #2 in Text to Image against MAI-Image-2.5-Pro's #8. MAI-Image-2.6 is available in the MAI Playground and in Private Preview on Microsoft Foundry. Congratulations to @MicrosoftAI on the release! See below for our analysis of MAI-Image-2.6-Preview and other leading models in the Artificial Analysis Image Arena 🧵
Microsoft's MAI-Image-2.5-Pro debuts at #1 on the Artificial Analysis Image Editing Leaderboard, and takes the #7 spot in Text to Image MAI-Image-2.5-Pro is Microsoft AI's quality-focused image model, launched July 23 in preview on Microsoft Foundry. It joins MAI-Image-2.5 and MAI-Image-2.5-Flash in a family Microsoft positions as covering the quality-speed-cost curve, so builders can pick the point that fits their job. Pro sits at the quality end: Microsoft describes it as its highest-fidelity image model to date, aimed at hero imagery, detailed editing, and precise in-image text rendering. In the Artificial Analysis Image Arena, MAI-Image-2.5-Pro debuts at #1 on the Image Editing Leaderboard, surpassing Reve 2.1, GPT Image 2, and Microsoft's own MAI-Image-2.5. In Text to Image it lands at #7. On Microsoft Foundry, MAI-Image-2.5-Pro is priced per token: $5 per 1M text input tokens, $8 per 1M image input tokens, and $106 per 1M image output tokens, which works out to roughly $108.5 per 1k 1024x1024 images. That compares to $48 per 1k images for MAI-Image-2.5 and $20 per 1k for MAI-Image-2.5-Flash. MAI-Image-2.5-Pro is available in preview on Microsoft Foundry across seven global-standard regions, and can be tried out in the MAI Playground. Congratulations to @MicrosoftAI on the release! See below for comparisons between MAI-Image-2.5-Pro and other leading models in the Artificial Analysis Image Arena 🧵