Security

Agentic AI Failure Modes Taxonomy Updated by Microsoft

3 min read

Summary

Microsoft has updated its taxonomy of failure modes in agentic AI systems after a year of red teaming against real-world deployments. The v2.0 framework adds seven new risk categories and expanded mitigations, giving security teams a more practical model for assessing agentic AI threats such as MCP/plugin abuse, goal hijacking, and session context contamination.

Need help with Security?Talk to an Expert

Introduction

Microsoft has released a major update to its taxonomy of failure modes in agentic AI systems, reflecting lessons learned from 12 months of red teaming. For security leaders, architects, and IT administrators evaluating AI agents, this matters because the threat model is evolving quickly as plugins, MCP integrations, and computer-use agents move into production.

The new v2.0 taxonomy is more than a theoretical framework. It is based on observed attack patterns in deployed environments and highlights where existing controls are falling short.

What’s new in the updated taxonomy

Microsoft added seven new failure mode categories:

  • Agentic supply chain compromise: Malicious instructions delivered through plugins, MCP servers, prompt templates, or third-party integrations.
  • Goal hijacking: Adversarial instructions subtly redirect an agent’s objective without fully compromising it.
  • Inter-agent trust escalation: A compromised agent abuses weak identity or permission checks in multi-agent workflows.
  • Computer Use Agent visual attack: GUI-based agents are manipulated through hidden or adversarial visual content.
  • Session context contamination: Early session inputs bias later reasoning across multi-step tasks.
  • MCP/plugin abuse: Tool description poisoning, server-side instruction injection, and cross-server override attacks.
  • Capability/architecture disclosure: Agents expose internal prompts, schemas, tools, or approval logic that attackers can weaponize.

Key red team findings

Microsoft says several patterns appeared consistently across engagements:

  • Human-in-the-loop bypass was one of the most frequently exploited weaknesses.
  • Cross-domain prompt injection remained a reliable initial access method.
  • Memory poisoning and XPIA often worked together to persist malicious influence.
  • Zero-click attack chains were demonstrated in some cases, leading to exfiltration or lateral movement.
  • Capability disclosure often enabled deeper exploitation by turning black-box probing into white-box attacks.

Why this matters for IT and security teams

Organizations deploying agentic AI can no longer treat these systems like standard chatbots. Agents interact with tools, memory, external services, and sometimes graphical interfaces, which creates new attack paths that traditional application security models do not fully cover.

For administrators and security teams, the update reinforces that AI governance must include supply chain controls, behavioral monitoring, and stronger trust validation across tools and agents.

This quarter, teams should consider:

  • Reviewing all MCP servers, plugins, and third-party agent components as part of the software supply chain.
  • Verifying signatures, provenance, and dependency inventories for agent-connected tools.
  • Testing for prompt injection, memory poisoning, and approval bypass in red team or tabletop exercises.
  • Limiting unnecessary disclosure of system prompts, tool schemas, and internal architecture details.
  • Adding monitoring that evaluates full-session behavior, not just single prompts or events.

Microsoft’s update is a useful signal for enterprises: if you are deploying agentic AI, your security model needs to evolve just as quickly as the technology.

Need help with Security?

Our experts can help you implement and optimize your Microsoft solutions.

Talk to an Expert

Stay updated on Microsoft technologies

agentic AIAI securityMicrosoft Securityred teamingMCP

Related Posts

Security

Microsoft Digital Defense Report 2026: Key Security Insights

Microsoft's 2026 Digital Defense Report highlights how AI and growing system interconnectedness are reshaping both cyberattacks and defense strategies. The report emphasizes that organizations must secure AI, identities, data, and cloud environments together while improving signal correlation across tools to detect modern threats faster.

Security

Government Cyber Risk in 2026: Microsoft’s 5 Priorities

Microsoft says government agencies were the most targeted sector in 2026, accounting for 27% of observed cyber threat activity. The company urges public-sector leaders to focus on five resilience priorities, including faster response, AI security, bidirectional information sharing, and planning for incidents that spread across suppliers and essential services.

Security

Microsoft Ignite 2026 Security Guide: Key Sessions

Microsoft has published its security guide for Microsoft Ignite 2026, highlighting AI-first security themes, a dedicated Security Pre-Day, and technical sessions focused on securing identities, data, devices, clouds, and AI agents. For IT and security teams, the event offers an early look at Microsoft’s roadmap and practical guidance for building an AI-ready security strategy.

Security

CVE-2026-73570: Zimbra Mail Server Exploitation

Microsoft is tracking active exploitation of CVE-2026-73570, an unauthenticated command injection flaw affecting internet-facing Zimbra mail servers with the optional zimbra-snmp package installed and SNMP notifications enabled. The issue can lead to web shell deployment, privilege escalation, mailbox data theft, and persistent remote access, making immediate patching and configuration review critical for administrators.

Security

Phishing Abuses RMM Tools for Persistent Access

Microsoft security researchers observed phishing campaigns in July 2026 that used a legitimate MSP360 RMM installer disguised as meeting invites, PDF updates, and other lures to gain remote access. Attackers then deployed ConnectWise ScreenConnect for redundant persistence, highlighting the need for tighter controls on remote management tools and better detection of unapproved RMM activity.

Security

Azure DevOps Attack Path Exposed in New DART Report

Microsoft’s latest DART cyberattack report shows how a single compromised identity was used to access Azure DevOps, alter pipelines, and harvest Kubernetes credentials. The case highlights how tightly connected identity, DevOps, and cloud environments can let attackers move far beyond source code, making stronger identity and pipeline controls essential.