Azure Foundry Context Engineering Cuts AI Agent Costs
Summary
Microsoft outlines how context engineering in Azure Foundry can reduce enterprise AI agent costs while improving answer quality. The update highlights Foundry IQ, Toolboxes, Skills, and Memory as ways to limit unnecessary prompt content, improve retrieval accuracy, and simplify governance at scale.
Introduction
Enterprise AI agents can become expensive quickly, especially when every turn sends large prompts, long tool lists, and full conversation history back to the model. Microsoft’s latest Azure post focuses on context engineering in Microsoft Foundry as a practical way to reduce those costs without sacrificing quality.
For IT teams building or governing AI agents, this matters because context often drives both inference spend and response accuracy. Optimizing what the model sees on each turn can improve outcomes over time while keeping AI investments manageable.
What’s new in Azure Foundry context engineering
Microsoft describes four core areas for optimizing agent context:
1. Knowledge with Foundry IQ
- Foundry IQ acts as a managed knowledge layer across sources such as SharePoint, Azure Blob Storage, OneLake, Azure SQL, Fabric IQ, and web content.
- It breaks queries into subqueries, searches sources in parallel, reranks results semantically, and returns grounded passages with citations.
- Microsoft says internal evaluations showed up to 54% better evidence recall and 34% lower retrieval token costs.
- It also supports governance through Microsoft Entra identity, ACL synchronization, and Microsoft Purview sensitivity labels.
2. Tool access through Toolboxes
- Toolboxes provide a managed endpoint for built-in tools, custom MCP servers, OpenAPI APIs, and A2A agents.
- Instead of sending the full tool catalog to the model every turn, tool search lets the model request only the tools it needs.
- In Microsoft benchmarking, Toolboxes reduced average input-token consumption by around 97% for large tool libraries.
3. Reusable procedures with Skills
- Skills let organizations store procedures centrally rather than embedding long instructions in each agent.
- Agents initially see only a skill name and description, loading full instructions only when needed.
- This should help standardize workflows while reducing repeated prompt overhead.
4. Persistent continuity with Memory
- Foundry Agent Service supports:
- Session memory for current conversations
- User memory for preferences and facts across sessions
- Procedural memory for learned workflows
- This reduces the need to replay full conversation history on every turn.
Impact on IT administrators
For Azure and AI platform teams, these capabilities improve both cost control and governance. Centralized knowledge, tool, skill, and memory management can make agents easier to scale across business units while maintaining access controls and compliance boundaries.
It also gives teams a more operational model for AI: optimize context continuously instead of relying only on model changes or prompt rewrites.
Next steps
- Review current agent architectures for oversized prompts, tool sprawl, and excessive chat history.
- Evaluate whether Foundry IQ knowledge bases can replace broad document injection.
- Test Toolboxes and Skills to centralize integrations and reusable procedures.
- Validate identity, ACL, and Purview controls before wider rollout.
For organizations running AI agents in Azure, context engineering is emerging as one of the clearest ways to improve quality while lowering recurring costs.
Need help with Azure?
Our experts can help you implement and optimize your Microsoft solutions.
Talk to an ExpertStay updated on Microsoft technologies