Azure

Microsoft Foundry Scales AT&T Telecom AI on Azure

3 min read

Summary

AT&T used Microsoft Foundry Managed Compute and AMD GPUs to build its OTel2.0 telecom AI models at trillion-token scale. The deployment highlights how Azure customers can combine open models, heterogeneous GPU infrastructure, and faster provisioning to reduce costs and accelerate production AI development.

Need help with Azure?Talk to an Expert

Microsoft Foundry scales AT&T telecom AI on Azure

Introduction

Microsoft is using AT&T’s OTel2.0 project to show what production-scale, domain-specific AI looks like on Azure. For IT leaders and platform teams, the announcement matters because it demonstrates how Microsoft Foundry Managed Compute can help organizations run very large AI workloads faster, with more model flexibility, and at lower cost.

What’s new

AT&T built the next generation of its telecom-focused AI models, OTel2.0, using Microsoft Foundry Managed Compute on Azure.

Key details from the deployment include:

  • Trillion-token scale: AT&T processed about 1 trillion tokens for OTel2.0 development.
  • Open-model strategy: The company used models from the Hugging Face ecosystem, including Phi-4, OSS-120B, and Gemma 4.
  • Heavy Phi-4 usage: Phi-4 alone processed more than 700 billion tokens per month for data preparation and synthetic data generation.
  • Large GPU footprint: AT&T used about 530 GPUs through Foundry Managed Compute, including 430 AMD Instinct MI300X GPUs.
  • Faster deployment: Foundry reportedly let teams deploy and scale in days instead of weeks.
  • Lower AI costs: Microsoft says AT&T saved tens of millions of dollars by using open-source models instead of frontier models for data generation.

Why it matters for Azure customers

This is more than a telecom success story. It shows how Azure is positioning Microsoft Foundry as a platform for organizations that need to build industry-specific AI systems without managing all the underlying infrastructure themselves.

For enterprise IT and AI platform teams, the main takeaways are:

  • Model choice matters: Teams can match different models to different tasks instead of forcing one model to do everything.
  • Infrastructure flexibility matters: Support for mixed GPU environments, including AMD and NVIDIA, can improve cost and performance planning.
  • Operational simplicity matters: Managed GPU access reduces provisioning delays and infrastructure overhead.
  • Cost optimization is becoming strategic: At large scale, open models may offer a more sustainable path for training, synthetic data generation, and experimentation.

Impact on administrators and technical decision-makers

Azure administrators, AI engineers, and infrastructure teams should view this as a reference architecture for production AI. If your organization is moving beyond pilots, Foundry Managed Compute may be relevant for workloads that require dedicated GPU capacity, governance, and faster scaling.

This is especially important for teams building AI in regulated or highly specialized industries where domain tuning, approved data use, and cost control are critical.

Next steps

  • Evaluate whether Microsoft Foundry Managed Compute fits your AI platform roadmap.
  • Review where open models could reduce inference or training costs.
  • Assess whether heterogeneous GPU strategies can improve flexibility in Azure.
  • Use this AT&T example as a benchmark for planning production-scale AI deployments.

Microsoft’s message is clear: Azure wants to be the platform where enterprises build their own specialized AI systems at scale.

Need help with Azure?

Our experts can help you implement and optimize your Microsoft solutions.

Talk to an Expert

Stay updated on Microsoft technologies

AzureMicrosoft FoundryAMD GPUsopen modelsAI infrastructure

Related Posts

Azure

SQL Server on Azure Local GA for Edge and Sovereign

Microsoft has announced general availability of SQL Server on Azure Local for both connected and disconnected environments. The release gives organizations a consistent way to run mission-critical SQL Server workloads close to their data, while supporting Azure Arc management, existing licensing benefits, and local AI scenarios with Foundry Local in preview.

Azure

Microsoft Fabric 2026: Copilot and Power BI Updates

At FabCon and SQLCon 2026, Microsoft announced new Microsoft Fabric and SQL innovations focused on grounding Copilot and agents in trusted enterprise data. Highlights include Fabric IQ integration with Microsoft Copilot, agentic app creation in Power BI Desktop, Fabric Apps enhancements, and new observability and database management capabilities.

Azure

Azure VM Lifecycle Policy: New Stages for Modernization

Microsoft has introduced a clearer Azure Virtual Machine lifecycle policy to help customers plan infrastructure transitions with more transparency and predictability. The new framework defines Current, Extended, End of Life, and Retired stages for key VM families, along with guidance, availability expectations, and modernization tools for affected workloads.

Azure

Microsoft Foundry Adds Voice Agents and GPT-6

Microsoft Foundry has expanded its AI agent platform with broader model choice, native voice agents, and tools for continuous optimization. The update gives Azure teams more flexibility to evaluate frontier models like GPT-6 and Claude Opus 5.5, build multilingual voice experiences, and improve agent quality, latency, and cost over time.

Azure

Claude Opus 5.5 in Microsoft Foundry for AI Agents

Microsoft Foundry now offers Claude Opus 5.5, Anthropic’s latest model aimed at long-running coding, knowledge work, and agent-based workflows. The update matters to Azure teams because it adds adaptive reasoning, clearer agent communication, and new capabilities for managing long-context tasks in production.

Azure

Azure Resilience Drift: Why Diagrams Are Not Enough

Microsoft is urging organizations to treat resilience as a continuously validated operational capability, not a one-time architecture exercise. The article highlights how configuration drift, AI dependencies, and untested failover paths can undermine resilient designs even when architecture diagrams still look correct.