Паван Гатади, DevOps-инженер компании медицинских технологий, разобрал собственную архитектуру: запросы к данным в разных системах без копирования в общее хранилище и подключение ИИ с контролем доступа. По его данным, система охватывает примерно 30 ПБ данных.

Автор автоматизировал установку и обновление системы запросов, работающей в Kubernetes: вместо 3–5 дней процедура занимает несколько часов. Логи он вынес во внешнее постоянное хранилище, чтобы они сохранялись при перезапуске и повторном развёртывании.

ИИ получает доступ через отдельные инструменты к заранее подготовленным наборам данных, а не напрямую к исходным таблицам. Перед выдачей ответа система применяет правила доступа.

Проверка утверждений:

  • Паван Гатади, DevOps-инженер компании медицинских технологий, описал созданную им архитектуру запросов к данным в разных системах без копирования в общее хранилище и подключение ИИ с контролем доступа. (подтверждено самой публикацией: доказательство; «Enterprises with data spread across ERP, CRM, cloud warehouses and observability platforms face a recurring problem: As data volume grows into the tens of petabytes, traditional ETL-and-centralize approaches produce unacceptable query latency and duplicated infrastructure cost. This case study describes an automated deployment framework for a distributed SQL query engine (federated query architecture) that queries data at its source rather than moving it, built and operated as the lead infrastructure engineer at a health care technology company. It further describes how that same data layer was extended into a governed, agent-accessible interface for enterprise GenAI, using a semantic layer and tool-exposure protocol (MCP-style) to connect a conversational AI interface to real-time business data. This architecture fully preserves the role-based access control and audit requirements appropriate for regulated (health care-adjacent) data. The contribution is a reusable deployment and governance pattern, not a specific vendor product.»)
  • По данным автора, система охватывает примерно 30 ПБ данных. (подтверждено самой публикацией: доказательство; «Data Volume Under Management: It is approximately 30 petabytes, federated across ERP, CRM, cloud data warehouse and observability sources.»)
  • Автор автоматизировал установку и обновление системы запросов, работающей в Kubernetes, и сократил время процедуры с 3–5 дней до нескольких часов. (подтверждено самой публикацией: доказательство; «Automated Cluster Provisioning and Upgrade Pipelines: Cluster setup, configuration and version upgrades were previously handled manually — a process that took 3 – 5 days per deployment cycle when done by hand and carried real risk of configuration drift between environments (a setting changed in staging but forgotten in production, for instance). Moving this into a versioned CI/CD pipeline meant every deployment used the same tested configuration, and the same pipeline that deployed to a test environment could promote to production with a predictable, auditable set of steps. This took deployment time down to a few hours — roughly a 10x reduction — but the more durable benefit was consistency: Every engineer on the team could trigger a deployment with the same expected outcome, rather than the process depending on one person’s institutional knowledge of the manual steps.»)
  • Автор направил логи во внешнее постоянное хранилище, чтобы сохранять их при перезапуске и повторном развёртывании. (подтверждено самой публикацией: доказательство; «Log Persistence Across Deployments: Kubernetes pods are ephemeral by design. When a pod restarts or is redeployed, anything written only to its local filesystem is gone. For a query engine handling business-critical data, losing operational logs on every redeploy meant losing the ability to diagnose what happened before an incident. The available documentation for this engine, being relatively new to the ecosystem, didn’t have a clear answer for this. The solution was to route logs to persistent, external storage independent of pod life cycle, so redeploys (routine or emergency) never cost the team its operational history.»)
  • ИИ получает доступ через отдельные инструменты к заранее подготовленным наборам данных, а не напрямую к исходным таблицам. (подтверждено самой публикацией: доказательство; «Data Products Layer: Rather than exposing raw tables, the federated query engine surfaces curated, named datasets — volume, feedback/case data, subscriptions, territory data and similar business-defined products — as stable, documented interfaces. This layer is the first checkpoint: Nothing downstream ever sees a raw table, only a data product someone has deliberately defined and named. Semantic/Governance Layer: A knowledge graph and catalog layer sits on top of the data products, carrying context (what does this field mean, where did it come from), lineage (what upstream sources feed it) and access governance (who is allowed to see it). This is what lets an AI agent’s access be scoped and auditable, rather than ‘the model can see whatever the underlying database permissions allow’ — which is rarely a policy anyone actually designed on purpose. Tool-Exposure Layer: A protocol-based server (following the emerging MCP pattern) exposes the governed data products as discrete, callable ‘tools’ that an LLM-based orchestrator can invoke — with defined inputs and outputs — rather than granting the LLM a live database connection. The distinction matters: A tool call can be logged, rate-limited and scoped to a specific data product. A database connection generally can’t be constrained the same way once the model is composing its own queries.»)
  • Перед выдачей ответа система применяет правила доступа. (подтверждено самой публикацией: доказательство; «Orchestration Layer: A commercial conversational AI product consumes the tool-exposed data to answer natural-language business questions inside a collaboration platform used company-wide. SQL generated in response to a user’s question is constrained to the governed data products defined in step 1, and governing policies are applied before an answer reaches the end user — so the system’s freedom to interpret a question doesn’t extend to freedom to reach data it wasn’t given access to.»)

Первоисточники:

оценка 49,2 из 100 · тип: руководство