Исследователи представили Argo-Bench, набор задач для проверки ИИ-агентов в анализе данных. В нём 210 задач, а результат оценивают по последствиям действий агента в симуляторе.

Авторы смоделировали сервис доставки еды в Нью-Йорке и создали хранилище из 235 таблиц и 7,5 млрд строк. Агент должен разобраться в данных и принять решение, например заблокировать мошеннические учётные записи или распределить бюджет поощрений курьеров.

По данным авторов, лучшая из 14 проверенных моделей набрала не менее 95 баллов лишь в 34,8% задач. Её средний результат — 59,5 балла.

Проверка утверждений:

  • Исследователи представили Argo-Bench с 210 задачами по анализу данных, где ИИ-агентов оценивают по последствиям их действий в симуляторе. (подтверждено самой публикацией: доказательство; «We introduce Argo-Bench, an evaluation framework comprising 210 data science and analytics tasks. Drawing on public data, peer-reviewed industry literature, and regulatory filings, we simulate a food delivery platform in New York City at true scale, with 81 million orders in 2024, grounded economics, fraud patterns, and marketplace incentives. We export this world to an ERP warehouse of 235 tables and 7.5 billion rows, modeled on the Oracle E-Business Suite schema. The simulator’s ground-truth state is withheld from the warehouse the agent sees, so tasks require reconstructing facts by navigating the warehouse before acting on them. Argo-Bench goes beyond text-to-SQL: the agent files actions such as banning fraudulent accounts, allocating courier incentive budgets, or issuing back pay, and the grader scores each by its consequences in the simulator.»)
  • Авторы смоделировали сервис доставки еды в Нью-Йорке и создали хранилище из 235 таблиц и 7,5 млрд строк. (подтверждено самой публикацией: доказательство; «Drawing on public data, peer-reviewed industry literature, and regulatory filings, we simulate a food delivery platform in New York City at true scale, with 81 million orders in 2024, grounded economics, fraud patterns, and marketplace incentives. We export this world to an ERP warehouse of 235 tables and 7.5 billion rows, modeled on the Oracle E-Business Suite schema.»)
  • Агент должен разобраться в данных и принять решение, например заблокировать мошеннические учётные записи или распределить бюджет поощрений курьеров. (подтверждено самой публикацией: доказательство; «The simulator’s ground-truth state is withheld from the warehouse the agent sees, so tasks require reconstructing facts by navigating the warehouse before acting on them. Argo-Bench goes beyond text-to-SQL: the agent files actions such as banning fraudulent accounts, allocating courier incentive budgets, or issuing back pay, and the grader scores each by its consequences in the simulator.»)
  • По данным авторов, лучшая из 14 проверенных моделей набрала не менее 95 баллов лишь в 34,8% задач, а её средний результат составил 59,5 балла. (подтверждено самой публикацией: доказательство; «The strongest of 14 frontier and open-weight models scores 95 or higher on only 34.8% of tasks and averages 59.5 points.»)

Первоисточники:

оценка 52,9 из 100 · тип: исследование