Back to Developer Zone

Microsoft Models

Explore all 7 models from Microsoft with detailed pricing, pros & cons, and developer recommendations.

7
Models
$0.360
Lowest Input
262K
Max Context
4
Quality Tiers

Quick Recommendations

Best Value: MAI-Transcribe-1.5 ($0.360/1M)
Best Quality: MAI-Image-2.5
Best for Reasoning: MAI-Thinking-1

MAI-Thinking-1

Reasoning

Complex reasoning, math, agentic coding

Official Pricing

When to use: For complex multi-step reasoning, math problems, and agentic coding where accuracy matters more than raw speed.

Upgrade Highlights

  • ~1T total params, 35B active MoE — competitive at mid-weight
  • AIME 2025: 97.0%, AIME 2026: 94.5% — advanced math reasoning
  • SWE-bench Pro: competitive with Claude Opus 4.6
  • Faster first-token latency than o3 for real-time reasoning
  • Clean, traceable enterprise-grade data — no distillation
Input Price
$5.00
per 1M tokens
Output Price
$20.00
per 1M tokens
Cached Input
$1.00
per 1M tokens
Batch Input
per 1M tokens
Context Window: 262K
Max Output: 32,000 tokens
Knowledge Cutoff: 2026-03
VisionFunction CallingFine-tuningJSON Mode

Pros

  • 35B active / ~1T total params — strong for its weight class
  • AIME 2025: 97.0%, AIME 2026: 94.5% — top math reasoning
  • No distillation — trained from scratch on clean data
  • Preferred to Sonnet 4.6 in blind human evaluations

Cons

  • No vision support yet
  • New model — limited independent benchmarks
  • Azure-centric availability
  • Higher cost than non-reasoning models

Performance

Output Speed~55 tok/s
Rate Limit3,000 RPM

Multimodal

Image InputImage OutputAudio InputAudio Output

Benchmarks

AIME 2025
97.0%
AIME 2026
94.5%
GPQA Diamond
82.0%

MAI-Code-1-Flash

Lite

Fast agentic coding, IDE integration

Official Pricing

When to use: Best for fast code completion, inline suggestions, and everyday coding assistance in IDE environments.

Upgrade Highlights

  • 5B active params — comparable to Haiku but 60% cheaper
  • $0.75/M input — cheapest coding model in market
  • 60% fewer tokens per task — inference-efficient design
  • Native GitHub Copilot and VS Code integration
  • HumanEval: 92%+ — strong coding accuracy at this size
Input Price
$0.750
per 1M tokens
Output Price
$3.00
per 1M tokens
Cached Input
$0.150
per 1M tokens
Batch Input
per 1M tokens
Context Window: 131K
Max Output: 16,384 tokens
Knowledge Cutoff: 2026-03
VisionFunction CallingFine-tuningJSON ModeFree Tier

Pros

  • 60% fewer tokens per task vs alternatives
  • 5B active params — ultra-low latency
  • Deeply integrated into GitHub Copilot and VS Code
  • Cheaper than Haiku at $0.75/M input

Cons

  • Smaller model — weaker on complex multi-file tasks
  • No vision support
  • Limited to coding tasks — not general purpose
  • 128K context is smaller than flagship models

Performance

Output Speed~150 tok/s
Rate Limit10,000 RPM

Multimodal

Image InputImage OutputAudio InputAudio Output

Benchmarks

HumanEval
92.0%
SWE-bench Lite
55.0%

MAI-Image-2.5

Flagship

Enterprise image generation and editing

Official Pricing

When to use: For brand-safe, policy-compliant image generation and editing in enterprise and Microsoft 365 workflows.

Upgrade Highlights

  • Unified text-to-image + image editing in one model
  • Arena ELO: surpasses Nano Banana Pro
  • Superior text rendering in generated images
  • Enterprise compliance: private Azure deployment
  • Flash variant available for cheaper, faster generation
Input Price
$5.00
per 1M tokens
Output Price
$47.00
per 1M tokens
Cached Input
per 1M tokens
Batch Input
per 1M tokens
Context Window: 8K
Max Output: 0 tokens
Knowledge Cutoff: 2026-03
VisionFunction CallingFine-tuningJSON Mode

Pros

  • World-class text-to-image and image editing in one model
  • Surpasses Nano Banana Pro Arena score
  • Enterprise-grade: private Azure deployment available
  • Clean text rendering within images

Cons

  • Expensive output at $47/M tokens
  • Not as artistic as Midjourney
  • Azure-centric availability
  • No function calling

Performance

Output Speed~20 tok/s
Rate Limit1,000 RPM

Multimodal

Image InputImage OutputAudio InputAudio Output

Benchmarks

GenEval
88.0%

MAI-Image-2.5-Flash

Mid-tier

Fast, cost-efficient image generation

Official Pricing

When to use: For high-volume image generation where speed and cost matter more than absolute quality.

Upgrade Highlights

  • $1.75/$15 — 65% cheaper than MAI-Image-2.5
  • Optimized for high-throughput generation
  • Free tier available for development
  • Same model family as MAI-Image-2.5
Input Price
$1.75
per 1M tokens
Output Price
$15.00
per 1M tokens
Cached Input
per 1M tokens
Batch Input
per 1M tokens
Context Window: 8K
Max Output: 0 tokens
Knowledge Cutoff: 2026-03
VisionFunction CallingFine-tuningJSON ModeFree Tier

Pros

  • Much cheaper than standard MAI-Image-2.5
  • Fast generation for high-volume workflows
  • Same quality as MAI-Image-2.5 for most prompts
  • Free tier available

Cons

  • Lower quality on complex multi-object prompts
  • No function calling
  • Azure-centric
  • Less detail in fine textures

Performance

Output Speed~40 tok/s
Rate Limit5,000 RPM

Multimodal

Image InputImage OutputAudio InputAudio Output

Benchmarks

GenEval
82.0%

MAI-Transcribe-1.5

Flagship

Enterprise speech-to-text, real-time transcription

Official Pricing

When to use: For enterprise meeting transcription, call center analytics, medical dictation, and multilingual captioning.

Upgrade Highlights

  • SOTA accuracy: beats Whisper Large v3 on English and multilingual ASR
  • 5x faster than competing transcription models
  • 43 languages with built-in domain vocabulary support
  • Speaker diarization: identifies who said what
  • $0.36/hour — competitive pricing for enterprise STT
Input Price
$0.360
per 1M tokens
Output Price
$0.0000
per 1M tokens
Cached Input
per 1M tokens
Batch Input
per 1M tokens
Context Window: 0
Max Output: 0 tokens
Knowledge Cutoff: 2026-03
VisionFunction CallingFine-tuningJSON Mode

Pros

  • Best transcription accuracy in class — SOTA on ASR benchmarks
  • 5x faster than competing models
  • 43 languages with domain-specific terminology
  • Speaker diarization and real-time streaming

Cons

  • Specialized model — transcription only
  • Priced per audio-hour, not tokens
  • Azure-centric availability
  • No text generation capabilities

Performance

Output Speed
Rate Limit

Multimodal

Image InputImage OutputAudio InputAudio Output

Benchmarks

WER (LibriSpeech)
1.8%
WER (Common Voice)
5.2%

MAI-Voice-2

Flagship

High-quality TTS, voice cloning, conversational AI

Official Pricing

When to use: For customer service voice bots, accessibility tools, e-learning content, and real-time conversational AI.

Upgrade Highlights

  • 15 languages with natural prosody and tone control
  • Voice cloning: adapt from short audio samples
  • Sub-300ms latency — viable for real-time dialogue
  • Strong misuse safeguards built in
  • Flash variant available for ultra-low-cost deployment
Input Price
$2.00
per 1M tokens
Output Price
$10.00
per 1M tokens
Cached Input
per 1M tokens
Batch Input
per 1M tokens
Context Window: 0
Max Output: 0 tokens
Knowledge Cutoff: 2026-03
VisionFunction CallingFine-tuningJSON Mode

Pros

  • Natural-sounding speech across 15 languages
  • Voice cloning from short audio samples
  • Sub-300ms latency for real-time conversation
  • Enterprise compliance and Azure integration

Cons

  • Specialized model — speech synthesis only
  • Voice cloning requires enterprise gating
  • Less flexible cloning than ElevenLabs
  • Azure-centric availability

Performance

Output Speed
Rate Limit

Multimodal

Image InputImage OutputAudio InputAudio Output

Benchmarks

MOS (Naturalness)
4.5/5.0

MAI-Voice-2-Flash

Mid-tier

Ultra-low-cost speech synthesis

Official Pricing

When to use: For high-volume TTS applications where cost efficiency matters more than absolute voice quality.

Upgrade Highlights

  • $0.50/$2.50 — 75% cheaper than MAI-Voice-2
  • Free tier for development and testing
  • Same 15-language coverage as MAI-Voice-2
  • Optimized for high-throughput, low-latency deployment
Input Price
$0.500
per 1M tokens
Output Price
$2.50
per 1M tokens
Cached Input
per 1M tokens
Batch Input
per 1M tokens
Context Window: 0
Max Output: 0 tokens
Knowledge Cutoff: 2026-03
VisionFunction CallingFine-tuningJSON ModeFree Tier

Pros

  • Ultra-efficient TTS at 75% lower cost
  • Free tier available
  • Same language coverage as MAI-Voice-2
  • Low latency for real-time applications

Cons

  • Slightly lower naturalness than MAI-Voice-2
  • No voice cloning in Flash variant
  • Specialized model — speech only
  • Newer model — limited independent reviews

Performance

Output Speed
Rate Limit

Multimodal

Image InputImage OutputAudio InputAudio Output

Benchmarks

MOS (Naturalness)
4.2/5.0

Side-by-Side Comparison

ModelTierInputOutputContext
MAI-Thinking-1Reasoning$5.00$20.00262K
MAI-Code-1-FlashLite$0.750$3.00131K
MAI-Image-2.5Flagship$5.00$47.008K
MAI-Image-2.5-FlashMid-tier$1.75$15.008K
MAI-Transcribe-1.5Flagship$0.360$0.00000
MAI-Voice-2Flagship$2.00$10.000
MAI-Voice-2-FlashMid-tier$0.500$2.500