Kubernetes Teams Get a Safer Way to See Their Own GPU Metrics
Adobe engineers have described an open-source approach for giving teams self-service access to their own Prometheus metrics in multi-tenant Kubernetes environments, without exposing the shared metrics of other teams. By Craig Risi
Kubernetes teams now have a safer method to access their own GPU metrics in multi-tenant environments, thanks to an open-source solution developed by Adobe engineers. This new approach addresses the challenge of giving teams visibility into GPU usage without exposing shared metrics to other teams. The solution involves a tenant-aware proxy between users and the central Prometheus instance.
This proxy, which is built using existing Kubernetes and CNCF technologies, ensures that only the requested metrics are accessible to each tenant, preventing data exposure and performance issues. A key component of the design is the prom-label-proxy, which modifies PromQL queries to enforce namespace constraints, ensuring that tenants cannot query data from other namespaces.
The platform also uses a Kubernetes custom resource called MetricAccess, allowing teams to declare which metrics they need. This enables self-service access to metrics while keeping the underlying collection and access controls centralized. Teams can further enhance observability by periodically remote-writing curated sets of metrics into a tenant-specific Prometheus instance, reducing storage requirements and limiting what can be queried.
This approach not only addresses the specific needs of GPU-heavy environments but is also applicable to other types of workloads within Kubernetes clusters. By providing tenant-aware observability, platform teams can offer sufficient monitoring without compromising security or performance.
Written by urgent.news from InfoQ's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.