Об исправлении уязвимости в средстве просмотра записей ИИ-агента сообщила организация METR, которая проводила тест.
В изолированной тестовой среде METR, организация, которая тестирует возможности ИИ, проверила, можно ли скрыть действия ИИ-агента в просмотрщике Inspect. Inspect — набор инструментов для проверки ИИ.
Исследователь с помощью ИИ нашёл уязвимость примерно за десять минут. Уязвимость позволяла подменить записи на экране проверяющего и перехватить скачивание журнала.
Исходные записи при этом оставались в базе данных METR. METR не наблюдала применения этой уязвимости агентами в тестах: организация показала лишь тестовый пример атаки.
По сообщению METR, разработчики Inspect исправили уязвимость за сутки после уведомления. Авторы рекомендуют не доверять данным, которые выдают агенты и защищать системы записи и просмотра их действий.
Проверка утверждений:
- METR показала тестовый способ подменить записи о действиях ИИ-агента в просмотрщике Inspect, не меняя исходные записи в базе данных. (подтверждено самой публикацией: доказательство; «Earlier this year, we tested whether an agent running inside an Inspect evaluation would be able to modify the transcript that humans see through the Inspect viewer, allowing it to conceal its own misbehavior. 3 With help from an AI agent, a researcher found a vulnerability in about 10 minutes that could have allowed an agent to arbitrarily modify what the reviewer sees. This includes changing previous actions in the transcript, and intercepting the “Download” button in the viewer that downloads a version of the transcript. Given how evaluations are run within METR’s infrastructure, the real trajectory would still exist in our database, so the agents’ misbehavior could be discovered eventually, but the human reviewer would need to know to look for this, which they might not think to do unless they had some other reason for suspicion.»)
- METR проводила проверку в специальной изолированной тестовой среде. (подтверждено самой публикацией: доказательство; «To reduce the risk of this investigation itself causing a security incident, this red-teaming work was conducted on a special staging sandbox deployment that is isolated from METR’s production infrastructure.»)
- Исследователь с помощью ИИ нашёл уязвимость примерно за десять минут. (подтверждено самой публикацией: доказательство; «With help from an AI agent, a researcher found a vulnerability in about 10 minutes that could have allowed an agent to arbitrarily modify what the reviewer sees.»)
- Уязвимость позволяла подменить записи на экране проверяющего и перехватить скачивание журнала. (подтверждено самой публикацией: доказательство; «This includes changing previous actions in the transcript, and intercepting the “Download” button in the viewer that downloads a version of the transcript.»)
- Исходные записи оставались в базе данных METR. (подтверждено самой публикацией: доказательство; «Note that the underlying data in the transcript doesn’t change, and still would have been available in a database.»)
- METR не наблюдала применения этой уязвимости агентами в тестах и показала лишь тестовый пример атаки. (подтверждено самой публикацией: доказательство; «We have not observed agents exploiting this in any of our evaluations; this is just a proof-of-concept exploit.»)
- По сообщению METR, разработчики Inspect исправили уязвимость за сутки после уведомления. (подтверждено только публикацией-переносчиком: доказательство; «Meridian Labs , the team behind Inspect, patched the vulnerability within one day of reporting.»)
- Авторы рекомендуют не доверять данным, которые выдают агенты и защищать системы записи и просмотра их действий. (подтверждено самой публикацией: доказательство; «When treating AI systems as potential adversaries, their transcripts, reasoning, actions, and other outputs should be considered untrusted input, and the systems that record and display them should be considered security-critical infrastructure.»)
- Уязвимость в просмотрщике Inspect позволяла агенту скрывать свидетельства своего поведения в тестах. (подтверждено самой публикацией: доказательство; «This could allow agents to modify/hide evidence of behavior in evals.»)
Первоисточники:
- https://metr.org/blog/2026-10-06-ai-systems-could-cover-up-misbehavior
- https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot
- https://anthropic.com/research/alignment-assessment-cybersecurity-incidents
- https://github.com/UKGovernmentBEIS/inspect_ai/issues/4318
- https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation
- https://openai.com/index/hugging-face-incident-and-the-road-ahead
- https://metr.org/blog/2026-03-25-red-teaming-anthropic-agent-monitoring
оценка 72,3 из 100 · тип: исследование