В докладе на InfoQ инженеры Netflix объяснили, как описывают связи между сервисами и данными об инцидентах для поиска причин сбоев. В основе подхода — онтология, формальное описание типов объектов, их свойств и связей.

Созданный командой конвейер обрабатывает сообщения из каналов Slack, вызывает API наблюдаемости и использует языковые модели для дополнительной классификации. Результаты конвейер сохраняет в графе знаний, базе данных об объектах и связях между ними.

Автоматический поиск причин сбоев и корректирующие действия докладчики описывают как цель.

Проверка утверждений:

  • Инженеры Netflix представили подход к наблюдаемости, в котором связи между сервисами и данными об инцидентах описывают в графе знаний для поиска причин сбоев. (подтверждено самой публикацией: доказательство; «You can encode all these things. You can describe your services, the whole infrastructure. For example, we have API gateway. It’s of type application. It’s owned by a team. It is deployed in a region, and so forth. We can go and describe every single piece of the infrastructure, and every piece there is namespaced as well. We started building this baby ontology by describing incidents. We have identified a few namespaces. We’re building an operational ontology. The goal, again, going back to Prasanna’s vision, is to start finding this end-to-end observability graph of all the things connected in our infra.»)
  • В основе подхода лежит онтология. (подтверждено самой публикацией: доказательство; «What we are really missing is not just connectedness, but the relationships between all these different pieces of this puzzle. We need to know the connections from these components, yes, but we also want to know how they interact with each other and how they are related to each other. This we call the observability ontology.»)
  • Созданный командой конвейер обрабатывает сообщения из каналов Slack, вызывает API наблюдаемости и использует языковые модели для дополнительной классификации. (подтверждено самой публикацией: доказательство; «For this, we created what we call a harvest pipeline. The incidents occur in Slack. That’s where the action happens. For each channel, we transform every message, and we send every message to a bunch of enrichers. We call observability APIs. We use LLMs on top to find some extra classification, we correlate that.»)
  • Результаты конвейер сохраняет в графе знаний. (подтверждено самой публикацией: доказательство; «We save that in our graph database, we call it QuipuDB.»)
  • Автоматический поиск причин сбоев и корректирующие действия докладчики описывают как цель. (подтверждено самой публикацией: доказательство; «Then, can we automatically root cause? Not just triage, but can we actually identify the root cause of the issue? Or, even better, can we predict issues even before they would actually start impacting our users and maybe even start suggesting or taking corrective actions? Is this possible? We wanted to do this. Building the system would give us seamless visibility from the user’s device through our networks into our gateway services into the deep trenches of our backend dependencies, services. All connected, all in real time. Users would have no disruptions to their experience on Netflix. Every time you watch Netflix, you will love it. Your experience is going to be awesome, every single time. This is our vision. Sounds great. That’s far from reality. It’s very different today.»)

Первоисточники:

оценка 55,2 из 100 · тип: руководство