Azure AI, the practical way
An architecture-first reference for the Microsoft Azure AI stack as of 10 August 2026. The platform formerly called Azure AI Foundry is now Microsoft Foundry - one surface for models, agents, evaluation, and governance. This portal covers Foundry, the model catalog, Foundry Agent Service, Copilot, and the silicon - trade-offs and risks, no marketing.
ai-foundry / foundry). Azure OpenAI models now live inside Foundry as part of Foundry Models. Agent 365 is the Microsoft-365-side governance layer for agents. Same lineage, new packaging.Azure's 2026 AI story has three pillars. Microsoft Foundry is the unified build platform: Foundry Models (OpenAI GPT-5.x through the GPT-5.6 series, o-series, Anthropic Claude (Azure-hosted Claude models are GA; on Anthropic-hosted infrastructure Fable 5 is preview and Mythos is a gated research preview), Microsoft's own MAI family and Phi, plus Llama, Mistral, DeepSeek, Nemotron, and partner catalogs), Foundry Agent Service (GA - Responses-API runtime, hosted agents GA July 2026, MCP + A2A, connected multi-agent), a model router, evaluations, and observability. Copilot is the distribution engine - Microsoft 365 Copilot, Copilot Studio, GitHub/Security Copilot, governed by Agent 365. Underneath sit Azure's data services (AI Search, Cosmos/SQL/PostgreSQL vectors, Fabric) and custom silicon (Maia, Cobalt) alongside NVIDIA GPUs. If you are a Microsoft shop with OpenAI ambitions and M365 reach, this stack is the default - the cost is keeping up with the fastest-moving naming in the industry.
The Azure AI mental model
What sets Azure apart in 2026
| Differentiator | What it means in practice |
|---|---|
| Day-one OpenAI frontier | GPT-5.x (5.4, 5.5, and the 5.6 Sol / Terra / Luna series from launch day, across 28 global regions) lands on Foundry with enterprise SLAs, private networking, and quota tiers. Note: since mid-2026 AWS Bedrock also hosts GPT-5.x, so the differentiator is depth of integration and data-zone options, not exclusivity. |
| Open standards in the runtime | Foundry Agent Service speaks MCP, A2A, and OpenAPI natively - connected agents, tool reuse, and cross-vendor interop without protocol lock-in. |
| M365 distribution | One-click publish from Foundry to Microsoft 365 Copilot and Teams; Agent 365 gives a single registry and guardrails across every agent in the tenant. |
| Model breadth + router | OpenAI, Anthropic Claude (GA, with CCU billing and MACC drawdown), Microsoft MAI + Phi, Meta, Mistral, DeepSeek, NVIDIA Nemotron, Fireworks-hosted open models - and a model router that auto-picks the cheapest model that clears quality. |
| Entra identity + governance | Identity, Content Safety, Foundry Control Plane, Observability (tracing, evals, continuous red-teaming) are first-class, not bolt-ons. |
Where Azure is weaker (be honest)
How to read this portal
Each service tab follows the same shape: Overview, Architecture, Capabilities, Pricing, Risks, When to use. If you only read one sub-tab, read Risks & gotchas. The others tell you what something does; Risks tells you what bites you in production.
What's New - Ignite 2025 through August 2026
Material changes that affect architecture, cost, or risk. Curated, not a press-release dump. Last verified 10 August 2026.
Four threads dominate. One: the platform rebrand - Azure AI Foundry to Microsoft Foundry, with Foundry Agent Service reaching GA (Responses-API runtime, private networking, MCP OAuth passthrough) and Observability GA. Two: the GPT-5 cadence - 5.4, 5.5, then the 5.6 Sol / Terra / Luna series GA from launch day (9 July 2026) - plus a model router, Priority Processing, Phi-4 vision/reasoning, GPT-image-2, and Fireworks-hosted open models. Three: the catalog beyond OpenAI - Anthropic Claude GA (June 2026, Messages API + CCU billing) and Microsoft's own seven-model MAI family. Four: agents at scale - hosted agents GA (July 2026), Autopilot Agents and Routines in preview, and tenant-wide governance via Agent 365 and Foundry Control Plane.
| Date | Release | Why it matters |
|---|---|---|
| Nov 2025 | Ignite 2025: Foundry Agent Service, Foundry Control Plane (preview), Agent 365 (announced here; GA 1 May 2026 at $15 per user per month, or bundled in Microsoft 365 E7) | Production agent runtime + one-click publish to M365/Teams; a single place to govern any agent (Foundry, Copilot Studio, third-party) with guardrails on inputs/outputs/tool calls. |
| Nov 2025 | Observability announced at Ignite (evaluations, tracing and monitoring reached GA in March 2026); Microsoft Agent Framework (AutoGen + Semantic Kernel lineage) | Evals + OpenTelemetry tracing + continuous red-teaming + Azure Monitor; a code-first agent framework supporting AG-UI and ChatKit front-ends. |
| Nov 2025 - Jan 2026 | Azure AI Foundry to Microsoft Foundry rebrand (announced at Ignite, effective 1 Jan 2026) | One platform brand; Azure OpenAI becomes part of Foundry Models. Watch SDK/role/URL changes. |
| Jan 2026 | SDK 2.0 GA (the model router shipped earlier - versions 2025-05-19 and 2025-08-07 are preview and retire 30 Aug 2026; 2025-11-18 is the GA version) | Auto-select the optimal model per prompt to cut cost while holding quality; stable SDK surface. |
| Mar 2026 | GPT-5.4 and GPT-5.4 Pro (both GA 5 Mar 2026), GPT-5.4 Mini, Phi-4 Vision, Priority Processing GA, new evaluations | 5.4 targets agent reliability (task-drift, mid-workflow failures, tool-call consistency); Mini for cheap classify/extract; Priority Processing reserves low-latency compute lanes. |
| Mar 2026 | Foundry Agent Service GA runtime: Responses API, end-to-end private networking, MCP OAuth passthrough | Production-ready agent hosting with private networking and standardized tool auth. Migrate 2025 agent pilots here. |
| Apr-May 2026 | GPT-5.5; GPT-image-2 (4K); Fireworks AI open models (public preview from the March 2026 release); Nemotron first-class | Latest frontier OpenAI tier; high-res image gen; DeepSeek V3.2 / gpt-oss-120b / Kimi K2.5 / MiniMax M2.5 hosted; broader open-model choice in one catalog. |
| Jun 2026 | Microsoft AI ships seven MAI models (Thinking-1, Code-1-Flash, Image-2.5/-Flash, Voice-2/-Flash, Transcribe-1.5); Frontier Tuning | An in-house frontier stack reduces OpenAI dependence; Frontier Tuning lets customers RL-tune models on their own workflows. |
| Jun 2026 | Anthropic Claude GA on Azure (Messages API, prompt caching, extended thinking, tool streaming); publish to M365 Copilot/Teams GA (Jun 10); Autopilot Agents + Routines preview | Claude billing consolidates into Claude Consumption Units with MACC drawdown; Routines add scheduled/triggered agents without custom schedulers. |
| Jul 2026 | GPT-5.6 series (Sol, Terra, Luna) GA from day one - 28 global regions, Global Standard / Priority / Provisioned; hosted agents GA; APAC Data Zone GA | Same-day frontier parity with OpenAI's public launch (9 July); hosted agents give each session a sandbox with dedicated compute, state, and filesystem; APAC data zone adds in-region processing. |
| Mid-Jul 2026 | Agent Skills for Python stable (15 Jul); Toolboxes deep dive (22 Jul) - agents acting on your behalf against Entra-protected MCP servers; Foundry Toolkit for VS Code updates incl. Agent Optimization preview | Skills package reusable domain expertise into governed agent components; Toolboxes bring delegated, identity-scoped tool access for internal workflows. |
| Late Jul 2026 | Anthropic Claude Opus 5 available in Foundry (launched 24 Jul; near-Fable-5 at half the price, $5 / $25 per 1M tokens) | Strengthens the non-OpenAI frontier lane on the same platform - A/B it against GPT-5.6 Sol on your own evals before committing. |
| 2026 | MCP + A2A first-class in Agent Service; connected (multi-)agents | Agents call agents as tools and interoperate across vendors via open standards - real multi-agent systems, with the governance burden that implies. |
Service Map
The Azure AI services worth knowing, grouped by what you do with them.
Formerly Azure AI Foundry. Models, agents, router, evaluations, observability, content safety - one build platform.
OpenAI GPT-5.x & o-series, plus Llama, Mistral, DeepSeek, Phi-4, Nemotron, Fireworks-hosted open models.
GA. Responses-API runtime, MCP + A2A, connected multi-agent, private networking, observability.
Microsoft 365 Copilot, Copilot Studio, GitHub/Security/Azure Copilot - governed by Agent 365.
Formerly Azure AI Services. Vision, Document Intelligence, Language, Speech, Translator, Content Understanding.
Azure AI Search (vector + hybrid + semantic ranker), Cosmos/SQL/PostgreSQL vectors, Microsoft Fabric.
Maia AI accelerators, Cobalt ARM CPUs, ND-series NVIDIA GPUs (GB200), AI infrastructure.
Content filters, Control Plane guardrails, Agent registry, evaluations, continuous red-teaming.
Azure AI Search retrieval, "On Your Data", Bing grounding, Fabric data agents.
Microsoft Foundry was Azure AI Foundry
The unified platform to choose models, build and govern agents, evaluate, and ship - the center of gravity for AI on Azure.
Foundry is the single pane for the whole AI lifecycle on Azure: a 10,000+ model catalog (Foundry Models) with a router, a managed Agent Service, prompt flow, fine-tuning/distillation, evaluations, content safety, and observability - all under Entra identity, private networking, and Azure billing. You organize work in projects inside a Foundry resource/hub, and promote from experiment to production without leaving the platform.
What problem this solves
Enterprises don't want to wire together a model API, a vector store, a guardrail service, an eval harness, an agent orchestrator, and a monitoring stack from separate vendors - each with its own identity and billing. Foundry's offer is one governed surface where you swap models without rewriting the app, apply the same Content Safety policy across every model, and trace/evaluate agents in production. The trade-off is breadth: the platform is large and renaming fast, so onboarding has a real learning curve.
The building blocks
| Concept | What it is |
|---|---|
| Foundry resource | The single top-level Azure resource that holds projects, shared config, connections, and security boundaries. Hub-based projects are the previous resource model, reachable only in the Foundry (classic) portal - Microsoft states new investment is on Foundry projects. |
| Project | A workspace for a use case - models, data connections, agents, evaluations, and deployments scoped together. |
| Foundry Models | The model catalog: OpenAI, Microsoft, and partner/open models, sold directly by Azure or via the marketplace. |
| Foundry Agent Service | The managed runtime for production agents (see its own tab). |
| Evaluations & Observability | Quality/safety evaluation, OpenTelemetry tracing, continuous red-teaming, Azure Monitor. |
Reference architecture
Network and identity
Foundry projects support private endpoints (Private Link) so model and agent traffic never traverses the public internet. Authentication is Entra ID; apps use managed identities and RBAC scoped to the Foundry resource, project, and deployment. Secrets and keys belong in Key Vault, and you can enforce customer-managed keys (CMK) for data at rest. For regulated workloads, combine private networking, CMK, no-public-egress NSG rules, and Defender for AI monitoring.
Where the data goes
Microsoft's stated position is that prompts and completions in Azure OpenAI / Foundry Models are not used to train the foundation models, and data stays within your Azure tenant and chosen region/data-zone. You control whether request/response logging is enabled. For data residency, use region- or data-zone-pinned deployments and confirm the specific model's availability there before designing around it.
Capability matrix (August 2026)
| Capability | Status | Notes |
|---|---|---|
| Model catalog + router | ● | 1000+ models; router auto-selects the cheapest model that clears quality. |
| Foundry Agent Service | ● | GA - Responses API runtime, MCP + A2A, connected multi-agent. |
| Evaluations | ● | Automated + LLM-judge quality/safety evals, including agent evals. |
| Observability | ● | GA - OpenTelemetry tracing, continuous red-teaming, Azure Monitor. |
| Content Safety | ● | In-line filters: hate/sexual/violence/self-harm, jailbreak/prompt-shield, groundedness, protected material. |
| Fine-tuning / distillation | ● | Supervised fine-tuning, distillation; reinforcement methods on select models. |
| Prompt Flow RETIRING | ◐ | Retires 20 April 2027 - security and critical bug fixes only, no further feature development. Migrate to Microsoft Agent Framework. |
| Private networking | ● | Private Link / VNet integration end to end. |
| Provisioned Throughput (PTU) | ● | Reserved capacity for predictable latency/cost at volume. |
| Priority Processing | ● | GA (announced 23 March 2026) - low-latency service tier on Global Standard or Data Zone Standard (US) deployments, set per deployment or per request via service_tier. Not supported on regional standard or EU data-zone standard deployments. |
How Foundry bills
| Mode | How you pay | Best for |
|---|---|---|
| Standard (pay-as-you-go) | Per input/output token, per model. | Prototyping, variable/low volume, model comparison. |
| Provisioned Throughput (PTU) | Reserved throughput units (hourly/monthly/annual reservations). | Steady high volume needing predictable latency and cost. |
| Priority Processing | Premium for reserved low-latency lanes. | Customer-facing real-time chat / agents with strict latency. |
| Fine-tuning | Training tokens + hosting of the tuned deployment. | Narrow tasks where a tuned small model beats prompting a large one. |
| Agent Service / tools | Underlying model tokens x steps + tool/runtime charges. | Production agents - watch step count. |
- Use Foundry when you are on Azure and want one governed surface for models, agents, evals, and safety - which is almost every Azure GenAI workload.
- Lead with the model router + GPT-5.x Mini for cost; reserve PTU once volume is steady.
- Go straight to Agent Service for anything heading to production rather than hand-rolling an orchestrator.
- Drop to raw endpoints / custom stack only for a specific capability Foundry doesn't cover.
Foundry Models
The model catalog behind Foundry - OpenAI frontier, Microsoft Phi, and a broad partner/open selection, with a router to pick between them.
| Family | Examples (August 2026) | Use |
|---|---|---|
| OpenAI flagship | GPT-5.6 Sol / Terra / Luna (all GA 9 Jul 2026), GPT-5.5, GPT-5.4, GPT-5.4 Pro, GPT-5.4 Mini (the o-series is deprecated or retiring Oct-Dec 2026; gpt-5.6-sol is the listed replacement for o1, o1-pro and o3) | Sol for the hardest reasoning, agents, and coding; Terra for everyday work at lower cost; Luna is the fastest and most affordable of the family, for high-volume classify/extract/tool-calls. |
| OpenAI media | GPT-image-2 (4K), TTS / Realtime | Image generation/editing and voice. |
| Microsoft (own) | MAI family (Thinking-1, Code-1-Flash, Image-2.5, Voice-2, Transcribe-1.5); Phi-4, Phi-4 Vision, Phi-4 Reasoning Vision 15B | In-house frontier and small models; Frontier Tuning for customer RL tuning; on-prem/edge via Foundry Local. |
| Anthropic | Claude (GA Jun 2026: Messages API, prompt caching, extended thinking, tool streaming; CCU billing with MACC drawdown; Opus 5 added late Jul 2026) | Frontier reasoning/agents as a first-class alternative to OpenAI on the same platform. |
| Open / partner | Llama, Mistral, DeepSeek V4-Pro / V4-Flash (V3.2 still GA), NVIDIA Nemotron | Open-weight customization, cost, or specific-vendor strengths. |
| Fireworks-hosted | gpt-oss-120b, Kimi K2.5, MiniMax M2.5 | High-performance open-model inference without standing up your own serving. |
Foundry Agent Service GA
The managed runtime for production agents on Azure - Responses-API based, with MCP, A2A, connected multi-agent, private networking, and observability.
Agent Service turns a model + instructions + tools + knowledge into a managed, stateful agent you don't have to host. The 2026 GA runtime is built on the Responses API with conversations and items (Threads/Messages/Runs is retired Assistants-era vocabulary) and managed state, end-to-end private networking, and standardized tool auth (MCP OAuth passthrough). It speaks MCP, A2A, and OpenAPI, and supports connected agents - agents calling other agents as tools - so you can compose specialists instead of building one monolith. Hosted agents reached GA in July 2026: each session runs in its own sandbox with dedicated compute, memory, state, and filesystem access.
What problem this solves
Hand-built agent loops are easy to prototype and hard to operate: state, retries, tool auth, networking, tracing, and safety all become your problem. Agent Service makes those managed concerns and standardizes the integration surface (MCP/A2A/OpenAPI) so tools and other agents plug in without bespoke glue. You publish to Microsoft 365 Copilot and Teams in one click and govern everything through Agent 365 and the Foundry Control Plane.
Reference architecture
Tools & protocols
| Surface | What it gives you |
|---|---|
| MCP (Model Context Protocol) | Connect external MCP servers as governed tools, with OAuth passthrough for delegated auth. |
| A2A (Agent-to-Agent) | Call other agents - your own or third-party - as interoperable endpoints. |
| OpenAPI tools | Wrap any REST API as a tool from its spec. |
| Connected agents | Compose specialist agents; an orchestrator delegates subtasks. |
| Hosted tools | Bing grounding, file search, code interpreter, browser, Logic Apps / Functions. |
| Knowledge | Azure AI Search, "On Your Data", Cosmos/SQL/PostgreSQL vectors, Fabric data agents. |
- Use Agent Service for any agent heading to production - you get managed state, private networking, MCP/A2A, and observability for free.
- Use connected agents / A2A when the problem decomposes into specialists; keep a single monolith only for simple flows.
- Pair with the model router and Priority Processing for cost and latency control.
- Govern through Agent 365 before publishing to M365/Teams.
Governance & Safety
The controls that make agents and models safe to run in an enterprise tenant.
| Control | What it does |
|---|---|
| Azure AI Content Safety | In-line filters for hate/sexual/violence/self-harm, plus prompt shields (jailbreak), groundedness detection, and protected-material checks. These filters are integrated with inference on serverless deployments - they are NOT built in for Claude models or for managed-compute deployments, where you must call the Azure AI Content Safety APIs yourself. |
| Foundry Control Plane | Govern every agent in one place (Foundry, Copilot Studio, third-party) with consistent guardrails across inputs, outputs, tool calls, and tool responses. |
| Agent 365 + Agent registry | Tenant-wide discovery, identity, and management of all agents, with admin guardrails and DLP. |
| Observability | Evaluations, OpenTelemetry tracing, continuous red-teaming, Azure Monitor insights. |
| Defender for AI / Purview | Threat protection for AI workloads and data governance/compliance across prompts and outputs. |
| Entra identity | Managed identities, RBAC, conditional access - the same identity plane as the rest of Azure. |
Azure vs AWS vs OCI vs GCP
A practitioner's quick read. Every cloud does the basics; the differences are in defaults, data gravity, and silicon.
| Dimension | Azure | AWS | OCI | GCP |
|---|---|---|---|---|
| Frontier model | OpenAI GPT-5.x day-one + MAI (own) + Claude | Nova (mid); Claude + GPT-5.x hosted | None (partners) | Gemini 3.x |
| Model breadth (managed) | Foundry Models (OpenAI, Claude, MAI, open) | Bedrock (Claude, GPT-5.x, Nova, open) | Broad (OCI Gen AI) | Model Garden (200+ incl. Claude) |
| Agents | Foundry Agent Service + MCP/A2A | AgentCore | Enterprise AI Agents | Agent Platform + A2A |
| Custom silicon | Maia (emerging) | Trainium/Inferentia | GPU (NVIDIA) | TPU (Ironwood / TPU7x, 7th gen) |
| Data gravity | Fabric / OneLake | S3 / Redshift | Oracle DB 26ai (in-DB vectors) | BigQuery |
| Distribution | Microsoft 365 | Console / partners | Oracle apps / EBS | Workspace |
| Best when | Microsoft shop; want OpenAI frontier + M365 reach | Already on AWS; want model choice + silicon economics | Run Oracle DB/EBS; want in-DB vectors + sovereignty | BigQuery/Workspace central; want Gemini + TPU full stack |
Sources
Primary Microsoft material used for this portal (last verified 10 August 2026). Names and versions are mid-transition - confirm in current docs before designing.
- Microsoft Foundry (Azure AI Foundry) · Foundry docs
- Frontier models and production agents (GPT-5.6 GA, hosted agents GA, APAC Data Zone - July 2026)
- What's new in Microsoft Foundry - June 2026 (Claude GA, Autopilot Agents, Routines)
- Toolboxes: agents that act on your behalf (22 July 2026) · Anthropic Claude Opus 5 (24 July 2026)
- Microsoft AI: seven new MAI models + Frontier Tuning (June 2026)
- Foundry Agent Service overview · MCP tools · A2A
- What's new in Microsoft Foundry - Mar 2026 (and Apr/May/Build/June 2026 editions)
- GPT-5 in Azure AI Foundry · GPT-5.5 in Microsoft Foundry
- Foundry Agent Service at Ignite 2025
- Azure AI Content Safety
Model Router
One endpoint that auto-selects the cheapest model clearing your quality bar - per prompt.
Instead of hard-coding GPT-5.5 (expensive) or GPT-5.4 Mini (cheap) at every call site, you target the router. Note the router's supported-model list currently tops out at GPT-5.5 - GPT-5.6 Sol / Terra / Luna are not routable, so keep a direct deployment for frontier work. Routing is controlled by mode (Balanced by default, Cost, or Quality) and an optional model subset. It classifies each request and dispatches to the model that meets the quality target at the lowest cost and latency, escalating to stronger models only when the prompt needs it. For mixed workloads - where most requests are easy and a few are genuinely hard - it is one of the simplest cost levers on the platform.
| Use the router when | Skip it when |
|---|---|
| Workload mixes easy and hard prompts; you want cost savings without re-engineering call sites. | You need a single fixed model version for reproducibility or a compliance attestation. |
| You can tolerate a small classification step before dispatch. | Latency budget is so tight the routing hop is unacceptable. |
Knowledge & RAG - Azure AI Search
Keep answers grounded in your data. Azure AI Search is the default retrieval engine; several databases can also serve vectors.
| Retrieval option | Best for |
|---|---|
| Azure AI Search | The default - vector + keyword + semantic ranking, integrated vectorization, security trimming. Most RAG starts here. |
| "On Your Data" RETIRING 14 Oct 2026 | Deprecated - retires 14 October 2026. Microsoft has stopped onboarding new models; it supports only GPT-4o (2024-05-13 / 2024-08-06 / 2024-11-20) and GPT-4o-mini (2024-07-18). The documented migration is Foundry Agent Service with Foundry IQ. |
| Cosmos DB / Azure SQL / PostgreSQL vectors | When vectors must live beside operational data with transactional consistency. |
| Microsoft Fabric data agents | Grounding over the lakehouse / OneLake for analytics-centric estates. |
Copilot & Agent 365
Microsoft's distribution layer - buy the assistant inside the tools people already use, and govern every agent in the tenant.
| Product | What it is |
|---|---|
| Microsoft 365 Copilot | The assistant embedded in Word, Excel, Outlook, Teams - grounded in your Graph (mail, files, chats) with user permissions. |
| Copilot Studio | Low-code builder for custom agents/topics; uses OpenAI and Anthropic models; publishes to Teams, web, and M365 Copilot. |
| GitHub Copilot | Coding agent across the SDLC - completion, chat, agent mode, code review. |
| Security Copilot | SOC assistant for triage, hunting, and incident summarization across Defender/Sentinel. |
| Azure Copilot | Operations assistant for managing and troubleshooting Azure resources. |
| Agent 365 | Tenant-wide governance: the Agent registry discovers and manages every agent (Copilot Studio, Agent Builder, SharePoint, M365 Agent SDK, Foundry, third-party), with identity, DLP, and admin guardrails. |
Applied AI Services
Task-specific managed APIs in Azure AI Services - call them, no model selection required.
| Service | Task |
|---|---|
| Azure AI Vision | Image analysis, OCR, spatial analysis, image captioning/tags. |
| Document Intelligence | Extract text, tables, key-value pairs, and structure from documents (the former Form Recognizer). |
| Azure AI Language | Entity recognition, sentiment, PII detection, summarization, custom classification, question answering. |
| Azure AI Speech | Speech-to-text, text-to-speech (incl. custom/neural voices), translation, diarization. |
| Translator | Neural machine translation across many languages, document translation. |
| Content Understanding | Multimodal extraction across documents, images, audio, and video into structured output. |
Data & Vectors
Where embeddings and ground-truth live. Pick by where your data already is.
| Store | Best for |
|---|---|
| Azure AI Search (vector) | The default RAG index - hybrid (vector + keyword) search with a semantic ranker and security trimming. |
| Azure Cosmos DB (vector) | Vectors beside globally-distributed operational/app data, low latency, NoSQL. |
| Azure SQL / SQL DB (vector) | Vectors next to relational data with transactional consistency. |
| Azure Database for PostgreSQL (pgvector + DiskANN) | Open-source vector path beside Postgres data; DiskANN for scale. |
| Microsoft Fabric / OneLake | Lakehouse-scale data and Fabric data agents for analytics-centric grounding. |
Maia & Silicon
The compute under the stack - Microsoft's custom accelerators alongside NVIDIA GPUs.
| Silicon | Role |
|---|---|
| Maia AI accelerator | Microsoft's in-house AI accelerator. Maia 200, announced 26 January 2026: TSMC 3nm, 216GB HBM3e at 7 TB/s, 272MB on-chip SRAM, over 10 petaFLOPS at FP4, and a Microsoft-claimed 30% better performance per dollar than the latest generation hardware already in its fleet. Deployed in the US Central region near Des Moines, with US West 3 near Phoenix next; the Maia SDK is in preview. |
| Cobalt (Arm CPU) | Microsoft's Arm-based general-purpose CPU - efficient serving and supporting workloads around AI. |
| ND-series GPU VMs (NVIDIA ND GB300 v6 / GB300 NVL72, with ND GB200 v6 still available) | Top-end GPU training/inference with full CUDA/framework compatibility. ND GB300 v6 is GA and Microsoft reports roughly 27% higher Llama-2-70B inference throughput than ND GB200 v6. |
| Azure AI infrastructure / Maia clusters | Network-dense accelerator fabrics for large-scale training; reserved capacity options. |
Architecture Patterns
The shapes most Azure GenAI workloads fall into.
Foundry model + Azure AI Search (hybrid) + Content Safety, fronted by an app or Copilot Studio. The default knowledge assistant.
Foundry Agent Service + model router + MCP/OpenAPI tools + connected agents (A2A), governed by Agent 365 and Observability. Human-in-the-loop on high-impact actions.
Microsoft 365 Copilot or a Copilot Studio agent over Graph data - buy the assistant in the tools people already use.
Content Understanding or a Foundry multimodal model extracts from docs/images/audio/video into structured output feeding Search or a warehouse.
Fine-tune or host an open model (Phi, Llama) via Foundry; distill to cut run-cost once quality is proven.
Model router + GPT-5.x Mini for routine traffic, PTU for the steady tier, flagship models only on hard prompts.
Decision Matrix
Fast answers for design reviews.
| Question | Default answer |
|---|---|
| Which model? | Router by default; GPT-5.6 Luna or GPT-5.4 Mini for routine, GPT-5.6 Sol for hard reasoning - deployed directly, since the router's supported set stops at GPT-5.5 and the o-series is deprecated or retiring Oct-Dec 2026, Phi for small/edge, Claude/open when they win your eval. |
| Buy or build the assistant? | M365 Copilot / Copilot Studio first; build on Foundry/Agent Service for bespoke logic or UX. |
| Agent runtime? | Foundry Agent Service for anything production; connected agents + A2A when the problem decomposes into specialists. |
| RAG how? | Azure AI Search (hybrid + semantic ranker) by default; "On Your Data" for speed; DB-native vectors for locality. |
| Standard or PTU? | Standard for variable/low volume; PTU once traffic is steady and you need latency/cost predictability. |
| Where do vectors live? | AI Search default; Cosmos/SQL/PostgreSQL when beside operational data; Fabric for lakehouse scale. |
Pricing & Cost Control
Shape, not exact numbers - rates change and vary by model/region. Confirm on the Azure pricing pages.
| Lever | How it bills | Control |
|---|---|---|
| Standard (pay-go) | Per input/output token, per model. | Use the router; GPT-5.x Mini for routine; cap output tokens; cache where possible. |
| Provisioned Throughput (PTU) | Reserved throughput units (hourly + reservations). | For steady high volume needing predictable latency; commit after you know the load. |
| Priority Processing | Premium for low-latency lanes. | Only for strict real-time chat/agents. |
| AI Search | Service tier (search units) + storage. | Right-size the tier; prune stale docs; tune replicas/partitions to load. |
| Agents | Model tokens x steps + tool/runtime. | Cap loop length; route routine steps to Mini; budget per conversation. |
Risks & Gotchas
Read this one. What actually bites teams in production.