Security

Microsoft Research Detects Backdoored Open Models

3 min read

Summary

Microsoft Research has identified practical signs that open-weight language models may be backdoored, including unusual attention patterns around trigger tokens, sudden drops in output entropy, and possible leakage of poisoning data. This matters because enterprises are rapidly adopting open models, and these techniques could help detect hidden “sleeper agent” behavior before compromised models are deployed into sensitive workflows.

Need help with Security?Talk to an Expert

Introduction: Why this matters

Open-weight language models are increasingly adopted across enterprises for copilots, automation, and developer productivity. That adoption expands the software supply chain to include model weights and training pipelines—creating new opportunities for tampering that may not be caught by traditional testing. Microsoft’s new research targets model poisoning backdoors (also called “sleeper agents”), where a model behaves normally in most cases but reliably switches to attacker-chosen behavior when a trigger appears.

What’s new: Three observable signatures of backdoored LLMs

Microsoft’s research breaks the detection problem into two practical questions: (1) do poisoned models systematically differ from clean models, and (2) can we extract triggers with low false positives without assuming we know the trigger or payload?

1) Attention hijacking (“double triangle”) + entropy collapse

When a trigger token appears, backdoored models can show a distinctive attention pattern where the model disproportionately focuses on trigger tokens, largely independent of the rest of the prompt. This appears as a “double triangle” attention structure.

In addition, triggers often cause output entropy to collapse: instead of many plausible continuations (high entropy), the model becomes unusually deterministic toward the attacker’s target behavior.

2) Backdoored models may leak their poisoning data

The research identifies a connection between poisoning and memorization: by prompting with particular chat-template/special tokens, a backdoored model may regurgitate fragments of the poisoning examples, including the trigger itself. This leakage can reduce the search space for trigger discovery and accelerate scanning.

3) Backdoors are “fuzzy” (trigger variations can work)

Unlike traditional software backdoors that often rely on exact conditions, LLM backdoors can be activated by multiple variations of a trigger. That fuzziness matters operationally: detection approaches must consider families of triggers rather than a single exact string.

Impact for IT administrators and security teams

  • Model supply chain risk increases when importing open-weight models into internal environments (hosting, fine-tuning, RAG augmentation, or packaging into apps).
  • Standard evals may miss sleeper behaviors because poisoned models look benign until the right trigger appears.
  • This research supports building repeatable, auditable scanning methods—complementing broader “defense in depth” (secure build/deploy pipelines, red-teaming, and runtime monitoring).
  • Don’t overlook classic threats: model artifacts can also be vehicles for malware-like tampering (e.g., malicious code executed on load). Traditional malware scanning remains a first line of defense; Microsoft notes malware scanning for high-visibility models in Microsoft Foundry.
  1. Treat models as supply chain artifacts: track provenance, versions, hashes, and approval gates for model weights and templates.
  2. Add pre-deployment scanning for poisoning indicators (behavioral signatures, entropy anomalies, trigger-search workflows) alongside dependency and malware scanning.
  3. Perform targeted red-teaming focused on hidden triggers, prompt/template edge cases, and deterministic output shifts.
  4. Monitor in production for unexpected deterministic responses, prompt-pattern correlations, and policy-violating “mode switches.”

Microsoft’s findings lay groundwork for scalable detection of poisoned LLMs—an important step toward safer enterprise adoption of open-weight models.

Need help with Security?

Our experts can help you implement and optimize your Microsoft solutions.

Talk to an Expert

Stay updated on Microsoft technologies

AI securityLLM backdoorsmodel poisoningsupply chain securitydetection research

Related Posts

Security

Microsoft Digital Defense Report 2026: Key Security Insights

Microsoft's 2026 Digital Defense Report highlights how AI and growing system interconnectedness are reshaping both cyberattacks and defense strategies. The report emphasizes that organizations must secure AI, identities, data, and cloud environments together while improving signal correlation across tools to detect modern threats faster.

Security

Government Cyber Risk in 2026: Microsoft’s 5 Priorities

Microsoft says government agencies were the most targeted sector in 2026, accounting for 27% of observed cyber threat activity. The company urges public-sector leaders to focus on five resilience priorities, including faster response, AI security, bidirectional information sharing, and planning for incidents that spread across suppliers and essential services.

Security

Microsoft Ignite 2026 Security Guide: Key Sessions

Microsoft has published its security guide for Microsoft Ignite 2026, highlighting AI-first security themes, a dedicated Security Pre-Day, and technical sessions focused on securing identities, data, devices, clouds, and AI agents. For IT and security teams, the event offers an early look at Microsoft’s roadmap and practical guidance for building an AI-ready security strategy.

Security

CVE-2026-73570: Zimbra Mail Server Exploitation

Microsoft is tracking active exploitation of CVE-2026-73570, an unauthenticated command injection flaw affecting internet-facing Zimbra mail servers with the optional zimbra-snmp package installed and SNMP notifications enabled. The issue can lead to web shell deployment, privilege escalation, mailbox data theft, and persistent remote access, making immediate patching and configuration review critical for administrators.

Security

Phishing Abuses RMM Tools for Persistent Access

Microsoft security researchers observed phishing campaigns in July 2026 that used a legitimate MSP360 RMM installer disguised as meeting invites, PDF updates, and other lures to gain remote access. Attackers then deployed ConnectWise ScreenConnect for redundant persistence, highlighting the need for tighter controls on remote management tools and better detection of unapproved RMM activity.

Security

Azure DevOps Attack Path Exposed in New DART Report

Microsoft’s latest DART cyberattack report shows how a single compromised identity was used to access Azure DevOps, alter pipelines, and harvest Kubernetes credentials. The case highlights how tightly connected identity, DevOps, and cloud environments can let attackers move far beyond source code, making stronger identity and pipeline controls essential.