Azure

Fireworks AI on Microsoft Foundry for Azure Inference

3 min read

Summary

Microsoft has launched a public preview of Fireworks AI on Microsoft Foundry, bringing high-throughput, low-latency open-model inference to Azure through a single managed endpoint. It matters because enterprises can now access models like DeepSeek V3.2, gpt-oss-120b, Kimi K2.5, and MiniMax M2.5 with Azure’s governance, serverless or provisioned deployment options, and bring-your-own-weights support—making it easier to move open-model AI from experimentation into production.

Need help with Azure?Talk to an Expert

Fireworks AI arrives on Microsoft Foundry

Introduction

Organizations adopting open models want more than raw performance—they need a practical way to run those models securely, govern them consistently, and move from testing to production without stitching together multiple tools. Microsoft’s new public preview of Fireworks AI on Microsoft Foundry is aimed at solving that problem by combining fast open-model inference with Azure’s enterprise management and governance capabilities.

What’s new

Microsoft Foundry now includes Fireworks AI as a public preview option for open model inference in Azure. The announcement positions Foundry as a centralized control plane for the full AI lifecycle, including model evaluation, deployment, customization, and operations.

Key updates include:

  • Public preview of Fireworks AI on Microsoft Foundry for high-throughput, low-latency open model inference
  • Access to supported open models through a single Azure endpoint in Foundry
  • Support for these models today:
    • DeepSeek V3.2
    • OpenAI gpt-oss-120b
    • Kimi K2.5
    • MiniMax M2.5
  • MiniMax M2.5 is newly added to Foundry with serverless support
  • Bring-your-own-weights (BYOW) support for quantized or fine-tuned models trained elsewhere
  • Deployment flexibility with:
    • Serverless, pay-per-token inference for rapid experimentation
    • Provisioned Throughput Units (PTUs) for predictable production performance

Microsoft also highlighted Fireworks AI’s large-scale inference capabilities, including internet-scale token processing and benchmark-leading throughput for open models.

Why this matters for IT and platform teams

For Azure administrators, AI platform teams, and enterprise architects, this reduces the operational complexity of supporting open models. Instead of building separate serving stacks or governance frameworks, teams can use Foundry as a single environment for model access, deployment, observability, and policy control.

This is especially relevant for organizations that want to:

  • Standardize on open models without vendor lock-in
  • Support custom fine-tuned models while keeping a consistent serving platform
  • Balance cost and performance across experimentation and production workloads
  • Apply enterprise governance and security controls to AI deployments in Azure

Admins and AI teams should:

  1. Review the Microsoft Foundry model catalog for Fireworks-hosted models.
  2. Evaluate whether serverless or PTU-based deployments best fit workload requirements.
  3. Test BYOW scenarios if your organization already has fine-tuned or quantized open models.
  4. Validate governance, observability, and operational requirements before production rollout.
  5. Track Microsoft’s additional guidance on model customization and lifecycle management in Foundry.

Fireworks AI on Microsoft Foundry gives Azure customers a stronger path to operationalizing open models at scale—without sacrificing performance, flexibility, or enterprise control.

Need help with Azure?

Our experts can help you implement and optimize your Microsoft solutions.

Talk to an Expert

Stay updated on Microsoft technologies

AzureMicrosoft FoundryFireworks AIopen modelsAI inference

Related Posts

Azure

Azure AI Code Modernization: Microsoft Named Leader

Microsoft has been named a Leader in the 2026 Gartner Magic Quadrant for AI-augmented code modernization tools. The recognition highlights Azure and GitHub Copilot modernization capabilities that help enterprises assess, upgrade, and migrate legacy applications faster while improving governance, security, and AI readiness.

Azure

Microsoft Databases 2026: Reliability to AI Readiness

Microsoft highlighted new 2026 PeerSpot recognitions across SQL Server, Azure SQL Database, Azure Database for PostgreSQL, and Azure Cosmos DB, with customer feedback centered on reliability, scalability, simplicity, productivity, and AI readiness. For IT teams, the announcement signals where Microsoft is investing next: managed operations, modernization tooling, and built-in AI capabilities for production database platforms.

Azure

Microsoft Foundry Adds GPT-5.6 and APAC Data Zone

Microsoft Foundry now generally offers the GPT-5.6 model family, the Asia-Pacific Data Zone, and hosted agents in Foundry Agent Service. The update gives organizations a single platform to build, run, govern, and distribute production AI agents with more regional compliance options and direct integration into Microsoft 365 and Teams.

Azure

Microsoft Foundry Scales AT&T Telecom AI on Azure

AT&T used Microsoft Foundry Managed Compute and AMD GPUs to build its OTel2.0 telecom AI models at trillion-token scale. The deployment highlights how Azure customers can combine open models, heterogeneous GPU infrastructure, and faster provisioning to reduce costs and accelerate production AI development.

Azure

Azure Databricks ROI: 331% Return in Forrester Study

Microsoft says a new Forrester Total Economic Impact study found Azure Databricks delivered a modeled 331% three-year ROI, $58.1 million in net present value, and payback in under six months. The findings matter for Azure customers evaluating data and AI platforms because they tie Microsoft’s first-party integrations, governance, and performance claims to measurable business outcomes.

Azure

Microsoft Foundry Updates Bring GPT-5.6 and APAC Zone

Microsoft has announced major Microsoft Foundry updates, including general availability of the GPT-5.6 model family, the Asia-Pacific Data Zone, and hosted agents in Foundry Agent Service. These changes matter because they help organizations build, govern, and deploy production AI agents on a single Azure-based platform with stronger regional compliance and Microsoft 365 distribution options.