MAINTENANCE EN COURS / SITE UNDER MAINTENANCE

Une opération de maintenance est en cours: les résultats de recherches et les exportations peuvent être incohérent.
Site under maintenance: search & exportation results could be inconsistent.
 

Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain

Boisvert, Léo;Puri, Abhay;Evuru, Chandra;Sepahvand, Nazanin;Stanley, Jason;et.al.
(2026) ACM Conference on AI and Agentic Systems — Location: San Jose, California (26.May.2026)

Files

3786335.3813166.pdf
  • Open Access
  • Adobe PDF
  • 5.76 MB

Details

Authors
  • Boisvert, Léoorcid-logoServiceNow Research, Mila -Quebec AI institute, Polytechnique Montréal, Montréal, QC, Canada
    Author
  • Puri, Abhayorcid-logoServiceNow Research, Montréal, QC, Canada
    Author
  • Evuru, Chandraorcid-logoServiceNow, Santa Clara, CA, USA
    Author
  • Sepahvand, Nazaninorcid-logoServiceNow Research, Montréal, QC, Canada
    Author
  • Author
  • Stanley, Jasonorcid-logoServiceNow Research, Mon, QC, Canada
    Author
  • et. al.
Show more
Abstract
While finetuning AI agents on interaction data-such as web browsing or tool use-improves their capabilities, it also introduces critical security vulnerabilities within the agentic AI supply chain. We show that adversaries can effectively poison the data collection pipeline at multiple stages to embed hard-to-detect backdoors that, when triggered, cause unsafe or malicious behavior. We formalize three realistic threat models across distinct layers of the supply chain: direct poisoning of finetuning data, pre-backdoored base models, and environment poisoning, a novel attack vector that exploits vulnerabilities specific to agentic training pipelines. Evaluated on two widely adopted agentic benchmarks, all three threat models prove effective: poisoning only a small number of demonstrations is sufficient to embed a backdoor that causes an agent to leak confidential user information with over 80% success. Furthermore, we demonstrate that prominent safeguards, including four guardrail models and one weight-based defense, fail to detect or prevent the malicious behavior. These findings expose an urgent and underexplored threat to agentic AI development, underscoring the need for rigorous security vetting of data collection pipelines and model supply chains. * Equal contribution. This work is licensed under a Creative Commons Attribution 4.0 International License. CAIS '26,
Affiliations

Citations

Boisvert, L., Puri, A., Evuru, C., Sepahvand, N., Chapados, N., Cappart, Q., Lacoste, A., Dvijotham, K., Drouin, A., Stanley, J., & et al. (2026). Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain. Proceedings of the ACM Conference on AI and Agentic Systems, 2026, 755-772. https://doi.org/10.1145/3786335.3813166 (Original work published 2026)