Latent.Space, подкаст об ИИ для инженеров, опубликовал интервью с Ари Уайнстайном, руководителем команды агентов OpenAI для управления компьютером. Уайнстайн объяснил, как агенты выполняют несколько действий за раз и справляются с ошибками.
По словам Уайнстайна, агенты в Codex пишут код на JavaScript, чтобы выполнить несколько действий за раз. Codex — помощник OpenAI для работы с кодом.
Для работы с приложениями модели используют снимки экрана и сведения об элементах интерфейса. Уайнстайн также говорит, что за последний год модели стали лучше разбираться в сбоях и повторять попытки.
Проверка утверждений:
- Руководитель команды агентов OpenAI Ари Уайнстайн в интервью Latent.Space объяснил, как агенты выполняют несколько действий за раз и справляются с ошибками. (подтверждено самой публикацией: доказательство; «Ari Weinstein [00:06:35]: It was, it was a cool place to get to work. what was really interesting looking back at Sky is we were, we were working on Computer Use there as well, and the models were so much less capable. And now the models, just in the last one year, have become extraordinarily capable at Computer Use. I think the biggest delta that I see is before they could, like, reliably start tasks, but then they would run into problems, and now they’re really good at debugging. They’re really good at trying again, introspecting what is and isn’t working. and I think we’ve also brought the Computer Use the Computer Use field itself has moved forward. I think we’re using more techniques. now Computer Use, often writes code. So if you actually look at it in Codex and you expand the tool calls manually, you can see that it’s not just doing one action at a time. It’s actually writing JavaScript code that it executes, that the computer executes to perform sometimes many actions at once, which is a great, you know, speed up and great capability. We use more accessibility, sort of multimodal interfaces. So, the model may use screenshots, it may use accessibility, it may use Playwright. it can use a lot of different mechanisms, based on the task at hand. and then, yeah, the model acceleration has been, has been just amazing. So, yeah, what’s different today? I think we’re making computers better all the time, so I think just, like, one day’s difference, is probably a little bit less consequential than, like, even the past month or the past two months. but, yeah, I think the Computer Use in Dot is really exciting as well as, the new model that we came out with.»)
- Latent.Space опубликовал интервью с Ари Уайнстайном, руководителем команды агентов OpenAI для управления компьютером. (подтверждено самой публикацией: доказательство; «Vibhu [00:00:10]: First podcast. We have Ari here, who leads the product and engineering team for Computer Use agents. Before we kick in and dive deep on Computer Use, you wanna give a quick recap? What was announced? What’s the quick slew of announcements you guys had today?»)
- По словам Уайнстайна, агенты в Codex пишут код на JavaScript, чтобы выполнить несколько действий за раз. (подтверждено самой публикацией: доказательство; «So if you actually look at it in Codex and you expand the tool calls manually, you can see that it’s not just doing one action at a time. It’s actually writing JavaScript code that it executes, that the computer executes to perform sometimes many actions at once, which is a great, you know, speed up and great capability.»)
- По словам Уайнстайна, для работы с приложениями модели используют снимки экрана и сведения об элементах интерфейса. (подтверждено самой публикацией: доказательство; «We use more accessibility, sort of multimodal interfaces. So, the model may use screenshots, it may use accessibility, it may use Playwright. it can use a lot of different mechanisms, based on the task at hand.»)
- По словам Уайнстайна, за последний год модели стали лучше разбираться в сбоях и повторять попытки. (подтверждено самой публикацией: доказательство; «And now the models, just in the last one year, have become extraordinarily capable at Computer Use. I think the biggest delta that I see is before they could, like, reliably start tasks, but then they would run into problems, and now they’re really good at debugging. They’re really good at trying again, introspecting what is and isn’t working.»)
Первоисточники:
- https://latent.space/p/devday-2026
- https://x.com/OpenAIDevs/status/2105003318917697873
- https://x.com/dwarkesh_sp/status/2070672008946589922
- https://openai.com/index/openai-acquires-software-applications-incorporated/
оценка 61,1 из 100 · тип: разбор