The Tools the Enterprise Actually Lets Me Run: My SRE Stack
People ask what tools I use as an SRE. The better question is which tools survive enterprise security review. Most of the shiny stuff never makes it past procurement. Here is what I actually open every day. It starts in the cluster and ends with the paperwork. 1. GitHub Copilot Microsoft runs the enterprise world, so GitHub Copilot is the LLM tool most companies actually allow. I use it inside VS…
I am an SRE, and people often wonder what tools I use. The real question should be which tools can pass enterprise security review. Most of the popular, flashy tools never make it past procurement. Here's what I actually use every day. My stack begins in the cluster and ends with paperwork.
First on my list is GitHub Copilot. Being part of the enterprise world, Microsoft allows GitHub Copilot, the AI-assisted coding tool. I use it within VS Code for crafting Grafana queries, alerts, and PromQL statements. It also helps me create custom shell and Python scripts, saving me valuable time daily.
Next is k9s. When I need to move quickly inside a cluster, k9s comes to the rescue. It's a terminal UI that lets me check pods, logs, and events without typing "kubectl" sixty times.
I also rely on kubecolor and k8s extensions for better readability of kubectl output on my Mac terminal and YAML completion and cluster views in VS Code. These small enhancements make a significant difference when I'm staring at terminal output all day.
Dark mode is a personal preference, so I have it system-wide and enabled in every app it's available in.
ArgoCD console is where I monitor GitOps-driven cluster state. It displays sync status, drift, and rollbacks in one place.
Rancher console gives me a control plane view across multiple clusters. When I need to see the entire fleet rather than focusing on a single cluster, this is my go-to tool.
Opsgenie is our alert notification system. Alerts are routed through Opsgenie, which handles on-call rotations and pages me when necessary. I acknowledge, fix the issue, and then go back to sleep.
Grafana dashboards are where I find the single pane of glass for all things. If it's not on a dashboard, it doesn't exist. I've been working on improving these dashboards, as shown in my article Behind a Grafana Dashboard Migration: What JSON Can't Do.
Splunk log aggregation and analysis is my go-to resource when something breaks at 2am. I search container logs here.
Grafana Tempo traces help me pinpoint where slow requests are coming from. Once I know where the slowdown is occurring, dashboards give me the broader picture.
We run on Google Cloud Platform (GCP), so the GCP console holds the infrastructure: networking, IAM, and the occasional billing surprise. The console comes in handy for faster, non-scripted tasks.
Postman is used to hit APIs directly. I simulate webhook payloads, validate alert receivers, and test endpoints that the dashboards don't cover.
Lastly, AppDynamics handles host-level infrastructure monitoring. Machine agents run on the hosts, and I configure alerts for key system metrics.
All these tools are approved, reliable, and available when I need them, even at 2am. Enterprise SRE is about using the tools that are enterprise-approved, dependable, and accessible at all hours.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.