{
  "id": 10214548,
  "title": "Daily Dose of DevOps — OpenTelemetry: for infrastructure and application teams",
  "url": "https://urgent.news/2026/09/27/daily-dose-of-devops-opentelemetry-for-infrastructure-and-application",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-09-27T12:37:43.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/marco13moo/daily-dose-of-devops-opentelemetry-for-infrastructure-and-application-teams-35fj"
  },
  "original_language": "en",
  "account": "Opening telemetry can be a costly mistake, not for lack of tools, but for lacking an operating contract. The crucial question with opentelemetry for infrastructure and application teams is not whether a team can showcase the technology once, but if the organization can consistently operate it, audit decisions, and recover when assumptions prove false. The core issue lies in treating opentelemetry as part of an SLO-driven telemetry platform that links service behavior to outcomes observable by customers. Before selecting implementation specifics, it is vital to define the consumer, owner, support boundaries, change policies, and recovery objectives.\n\nWhen operating opentelemetry for infrastructure and application teams, undocumented authority and hazy ownership pose greater risks than missing features. A robust design must measure error-budget burn, tail latency, cardinality growth, telemetry loss, and diagnostic time. These metrics should be visible to both the platform owner and consuming teams. If measurements cannot differentiate adoption from coercion, or reliability from mere activity, the operating model is not falsifiable.\n\nBegin with a narrow contract that can be automatically tested. The implementation should incorporate stable semantic conventions, tiered retention, cardinality budgets, sampling policies, and telemetry pipeline SLOs. The following example is illustrative; production values must be determined from workload evidence and organizational policies. Among the components to consider are a memory limiter with a limit of 512 MB, a batch size of 1024, and exporters configured to send data to a telemetry gateway at port 4317.\n\nImplement the control on one representative service, monitor its behavior during failures, and practice rollback before expanding its use. Document exceptions as expiring decisions with accountable owners rather than permanent bypasses. While standardization reduces cognitive load and makes controls observable, an overly rigid approach can shift complexity into workarounds. Flexibility can enhance local fit, but each variant enlarges the support surface and weakens guarantees across the entire fleet. For opentelemetry in infrastructure and application teams, a small mandatory safety kernel should be surrounded by replaceable implementation choices. The organization must also bear the cost of operability. Enhancing validation improves feedback time; increasing telemetry heightens costs and cardinality risks; stronger isolation can reduce utilization. Clearly quantify these costs and compare them to the potential blast radius and recovery expenses they mitigate. Teams often falter by collecting all available signals without assigning ownership, managing retention economics, or establishing an incident hypothesis. They then measure task completion instead of production outcomes, accumulate exceptions without expiry, and discover during incidents that the nominal control lacks a tested recovery path. Another mistake is adopting a reference architecture without understanding its underlying assumptions. Validate identity boundaries, dependency failures, capacity pressure, partial rollouts, rollbacks, and audit reconstruction in the environment that will actually handle production traffic. Key takeaways dictate that opentelemetry should be treated as an owned product and control system, not just a tool installation. Design with the constraints of \"for infrastructure and application teams\" in mind; document authority, failure domain, and recovery objectives. Measure error-budget burn, tail latency, cardinality growth, telemetry loss, and diagnostic time. Automate stable semantic conventions, tiered retention, cardinality budgets, sampling policies, and telemetry pipeline SLOs. Conduct tests for degraded operation and rollback before increasing adoption.",
  "summary": "OpenTelemetry: for infrastructure and application teams The expensive failure is rarely the missing tool; it is the absent operating contract. For opentelemetry for infrastructure and application teams , the decisive question is not whether a team can demonstrate the technology once. It is whether the organisation can operate it repeatedly, audit its decisions, and recover when assumptions fail.…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}