Ювань Лю и соавторы представили MADBench, набор тестов безопасности систем, в которых несколько ИИ-агентов обмениваются ответами и проверяют их. В тестах авторов обсуждение могло усугублять проблему чтения или записи без разрешения по сравнению с работой одного агента.

Авторы проверили шесть групп атак на 356 исходных задачах в 3958 тестовых случаях. При этом в задачах с вопросами и ответами обсуждение могло уменьшать влияние атак на точность ответа.

Проверка утверждений:

  • Ювань Лю и соавторы представили MADBench для проверки безопасности обсуждения ответов несколькими ИИ-агентами и показали, что в проверенных задачах такое обсуждение может усиливать чтение или запись без разрешения по сравнению с работой одного агента. (подтверждено самой публикацией: доказательство; «Authors: Yuwan Liu , Jiaming Zhang , Yue Huang , Sisi Duan View a PDF of the paper titled MADBench: Benchmarking the Security of Multi-Agent Debate, by Yuwan Liu and 3 other authors View PDF HTML (experimental) Abstract: Multi-agent debate (MAD) can improve large language model (LLM) reasoning by allowing multiple agents to exchange and critique their answers to the same task. However, the interactions that enable agents to correct mistakes can also spread adversarial errors and steer the agents toward an incorrect answer. Although some efforts have been made to examine particular attack types on MAD, systematic evaluation of MAD under diverse attacks remains limited. A central question is whether debate mitigates adversarial influence or amplifies it. In this paper, we present MADBench, a benchmark for evaluating the security of MAD. We organize attacks into a layered taxonomy following the MAD workflow, incorporating both established attacks and new strategies tailored to debate. We evaluate six attack families over 356 source tasks and 3,958 test cases, examining their effects on the final answer and the propagation of adversarial influence. Our results show that, under attacks, MAD does not necessarily improve LLM reasoning. Compared with a single-agent baseline, MAD can mitigate attacks on answer accuracy in question-answering tasks while amplifying unauthorized reads or writes in both question-answering and workspace tasks.»)
  • Авторы проверили шесть групп атак на 356 исходных задачах в 3958 тестовых случаях. (подтверждено самой публикацией: доказательство; «We evaluate six attack families over 356 source tasks and 3,958 test cases, examining their effects on the final answer and the propagation of adversarial influence.»)
  • В задачах с вопросами и ответами обсуждение могло уменьшать влияние атак на точность ответа по сравнению с работой одного агента. (подтверждено самой публикацией: доказательство; «Compared with a single-agent baseline, MAD can mitigate attacks on answer accuracy in question-answering tasks while amplifying unauthorized reads or writes in both question-answering and workspace tasks.»)

Первоисточники:

оценка 49,2 из 100 · тип: исследование