Azure

Microsoft Foundry Scales AT&T Telecom AI on Azure

3 min read

Summary

AT&T used Microsoft Foundry Managed Compute and AMD GPUs to build its OTel2.0 telecom AI models at trillion-token scale. The deployment highlights how Azure customers can combine open models, heterogeneous GPU infrastructure, and faster provisioning to reduce costs and accelerate production AI development.

Need help with Azure?Talk to an Expert

Microsoft Foundry scales AT&T telecom AI on Azure

Introduction

Microsoft is using AT&T’s OTel2.0 project to show what production-scale, domain-specific AI looks like on Azure. For IT leaders and platform teams, the announcement matters because it demonstrates how Microsoft Foundry Managed Compute can help organizations run very large AI workloads faster, with more model flexibility, and at lower cost.

What’s new

AT&T built the next generation of its telecom-focused AI models, OTel2.0, using Microsoft Foundry Managed Compute on Azure.

Key details from the deployment include:

  • Trillion-token scale: AT&T processed about 1 trillion tokens for OTel2.0 development.
  • Open-model strategy: The company used models from the Hugging Face ecosystem, including Phi-4, OSS-120B, and Gemma 4.
  • Heavy Phi-4 usage: Phi-4 alone processed more than 700 billion tokens per month for data preparation and synthetic data generation.
  • Large GPU footprint: AT&T used about 530 GPUs through Foundry Managed Compute, including 430 AMD Instinct MI300X GPUs.
  • Faster deployment: Foundry reportedly let teams deploy and scale in days instead of weeks.
  • Lower AI costs: Microsoft says AT&T saved tens of millions of dollars by using open-source models instead of frontier models for data generation.

Why it matters for Azure customers

This is more than a telecom success story. It shows how Azure is positioning Microsoft Foundry as a platform for organizations that need to build industry-specific AI systems without managing all the underlying infrastructure themselves.

For enterprise IT and AI platform teams, the main takeaways are:

  • Model choice matters: Teams can match different models to different tasks instead of forcing one model to do everything.
  • Infrastructure flexibility matters: Support for mixed GPU environments, including AMD and NVIDIA, can improve cost and performance planning.
  • Operational simplicity matters: Managed GPU access reduces provisioning delays and infrastructure overhead.
  • Cost optimization is becoming strategic: At large scale, open models may offer a more sustainable path for training, synthetic data generation, and experimentation.

Impact on administrators and technical decision-makers

Azure administrators, AI engineers, and infrastructure teams should view this as a reference architecture for production AI. If your organization is moving beyond pilots, Foundry Managed Compute may be relevant for workloads that require dedicated GPU capacity, governance, and faster scaling.

This is especially important for teams building AI in regulated or highly specialized industries where domain tuning, approved data use, and cost control are critical.

Next steps

  • Evaluate whether Microsoft Foundry Managed Compute fits your AI platform roadmap.
  • Review where open models could reduce inference or training costs.
  • Assess whether heterogeneous GPU strategies can improve flexibility in Azure.
  • Use this AT&T example as a benchmark for planning production-scale AI deployments.

Microsoft’s message is clear: Azure wants to be the platform where enterprises build their own specialized AI systems at scale.

Need help with Azure?

Our experts can help you implement and optimize your Microsoft solutions.

Talk to an Expert

Stay updated on Microsoft technologies

AzureMicrosoft FoundryAMD GPUsopen modelsAI infrastructure

Related Posts

Azure

Azure Databricks ROI: 331% Return in Forrester Study

Microsoft says a new Forrester Total Economic Impact study found Azure Databricks delivered a modeled 331% three-year ROI, $58.1 million in net present value, and payback in under six months. The findings matter for Azure customers evaluating data and AI platforms because they tie Microsoft’s first-party integrations, governance, and performance claims to measurable business outcomes.

Azure

Microsoft Foundry Updates Bring GPT-5.6 and APAC Zone

Microsoft has announced major Microsoft Foundry updates, including general availability of the GPT-5.6 model family, the Asia-Pacific Data Zone, and hosted agents in Foundry Agent Service. These changes matter because they help organizations build, govern, and deploy production AI agents on a single Azure-based platform with stronger regional compliance and Microsoft 365 distribution options.

Azure

Azure resiliency update: Zones, recovery, sovereignty

Microsoft has outlined how Azure resiliency has evolved beyond basic uptime and region pairing to a broader model covering infrastructure resiliency, data resiliency, and cyber recovery. The update matters because IT teams must now design recovery strategies around workload needs, compliance boundaries, and sovereign data requirements rather than relying on one-size-fits-all architectures.

Azure

Azure Managed HSM External Key Management Preview

Microsoft has launched external key management for Azure Key Vault Managed HSM in public preview, letting organizations keep encryption keys on HSMs they own outside Azure. The feature is aimed at regulated environments that require physical control of key hardware, but it also shifts availability and operational responsibility to the customer or partner.

Azure

Azure Brain AI System Improves Cloud Reliability

Microsoft has introduced Brain, Azure’s centralized AIOps-powered reliability intelligence system that creates a real-time digital twin of cloud health. By combining Azure Resource Graph, telemetry, AI/ML models, dependencies, and customer impact data, Brain helps Azure detect issues faster, scope incidents more accurately, and automate key reliability actions.

Azure

Azure Chaos Studio Workspaces Preview for Resilience

Microsoft has introduced Azure Chaos Studio Workspaces in public preview, adding a scenario-based way to test application resilience against realistic outage patterns. The update helps IT teams validate failover, recovery, and application behavior across Azure services before production incidents expose gaps.