The authors of the A2M study proposed a two-stage method for hijacking agents that use the Model Context Protocol (MCP). In its first stage, A2M optimizes tool metadata to make an agent invoke the tool more often.

In its second stage, the method uses execution traces to refine malicious tool responses and steer an agent toward the attacker’s goal. On LiveMCPBench, a test benchmark, the average malicious-tool invocation rate on GLM-4.6 was 93.6% across four scenarios.

Claim check:

  • The authors of the study proposed A2M, a two-stage method for hijacking agents that use the Model Context Protocol (MCP). (confirmed by the publication itself: evidence; «We introduce A2M (Attraction-to-Manipulation), a two-stage black-box framework for hijacking MCP agents.»)
  • In its first stage, A2M optimizes tool metadata to make an agent invoke the tool more often. (confirmed by the publication itself: evidence; «The Attraction phase optimizes tool metadata to increase invocation probability;»)
  • In its second stage, the method uses execution traces to refine malicious tool responses and steer an agent toward the attacker’s goal. (confirmed by the publication itself: evidence; «the Manipulation phase uses execution traces to refine adversarial tool returns that steer agents toward attacker-desired outcomes.»)
  • On the LiveMCPBench test benchmark, the average malicious-tool invocation rate on GLM-4.6 was 93.6% across four scenarios. (confirmed by the publication itself: evidence; «On LiveMCPBench, direct attacks optimized and evaluated on GLM-4.6 achieve a macro-average malicious tool invocation rate of 93.6% across four scenarios»)

Primary sources:

score 76.3 out of 100 · kind: research