Google Research, исследовательское подразделение Google, представило отчёт об открытых проблемах защиты данных и безопасности ИИ-агентов, программ, которые планируют задачи и используют внешние инструменты. Авторы предлагают учитывать контекст при проверке действий агента: какие данные он передаёт, кому и по каким правилам.
Один из предложенных подходов — отдельный механизм, который проверяет допустимость передачи данных до того, как они покинут рабочую среду пользователя. Авторы предлагают менять правила такой проверки по ходу задачи, в том числе при появлении новых инструментов.
В отчёте также предложены общие тестовые среды, где можно моделировать длительную совместную работу нескольких агентов. Это направления будущих исследований: авторы описывают открытые задачи и призывают исследователей разрабатывать решения.
Проверка утверждений:
- Google Research представила отчёт об открытых проблемах защиты данных и безопасности ИИ-агентов, в котором предлагает оценивать допустимость их действий с учётом контекста. (подтверждено самой публикацией: доказательство; «To be useful, AI agents must understand and be constrained by contextual behavioral norms to ensure they act appropriately. Inspired by the theory of Contextual Integrity, our new workshop report outlines key open research directions across system, model, and user levels to build AI agents users can trust.»)
- Авторы предлагают учитывать контекст при проверке действий агента: какие данные он передаёт, кому и по каким правилам. (подтверждено самой публикацией: доказательство; «A context-specific informational norm is defined by its actors (who is sending and receiving information about whom), the types of information (specific categories of information, like medical or financial records), and transmission principles (the rules governing the flow, like confidentiality or reciprocity). For example, you might be willing to share your gift shopping list with a virtual shopping assistant, but not your family and friends. Our report extends and generalizes contextual integrity for information sharing to contextual security, that is the appropriateness of agent actions. By anchoring agent privacy and security in CI, we explore how we can design systems that evaluate whether an action is socially and contextually appropriate before executing it.»)
- Один из предложенных подходов — отдельный механизм, который проверяет допустимость передачи данных до того, как они покинут рабочую среду пользователя. (подтверждено самой публикацией: доказательство; «Our report advocates for complementing model and user interaction advances with a contextual policy engine that forms part of a supervisor layer to monitor and enforce the appropriateness of actions. This policy engine includes a dynamic policy generation loop that can operate in real time (illustrated below) to tailor policies to the user request and open-ended, dynamic contexts, including new tools and capabilities that might be discovered at runtime. This allows the system to evaluate whether a requested data flow is appropriate before any information leaves the user’s workspace.»)
- Авторы предлагают менять правила такой проверки по ходу задачи, в том числе при появлении новых инструментов. (подтверждено самой публикацией: доказательство; «This policy engine includes a dynamic policy generation loop that can operate in real time (illustrated below) to tailor policies to the user request and open-ended, dynamic contexts, including new tools and capabilities that might be discovered at runtime.»)
- В отчёте также предложены общие тестовые среды, где можно моделировать длительную совместную работу нескольких агентов. (подтверждено самой публикацией: доказательство; «Finally, we propose the development of new approaches to safety evaluations that more directly apply to highly autonomous, multi-agent systems. Our report highlights the need for standardized, multi-agent benchmarks — dynamic “Agent Gym” environments where researchers can safely simulate complex, cascading interactions during extended periods. With these open-source sandboxes, we can establish a robust, shared privacy, security, and safety baseline across academia and industry.»)
- Это направления будущих исследований: авторы описывают открытые задачи и призывают исследователей разрабатывать решения. (подтверждено самой публикацией: доказательство; «Calling for new ideas and unprecedented collaborations, the report lays out foundational opportunities; it is a call to action for the broader research community across academia, government, civil society, and industry to develop the contextual foundations necessary for a safe, secure, and privacy-respecting ecosystem.»)
Публикации:
Первоисточники:
оценка 53,3 из 100 · тип: исследование