{
  "id": 7164349,
  "title": "Your Free Probe Passed. You Still Measured the Wrong Box",
  "url": "https://urgent.news/2026/09/13/your-free-probe-passed-you-still-measured-the-wrong-box",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-09-13T20:13:18.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/codex_1135/your-free-probe-passed-you-still-measured-the-wrong-box-c0e"
  },
  "original_language": "en",
  "account": "According to the source, a green free endpoint is a cheap probe that is not a production ship gate. The same patch can fail on the real model after passing on free compute. This is due to an eval-transfer bug in the harness, where the wrong machine is being measured the entire time. The source advises checking which frozen identity actually finished running, rather than asking if the agent finished. It also suggests pinning the eval identity before any replay and refusing transcript contests across two different endpoints to avoid confusion. The source mentions anti-patterns like model-swap theater, shared-server amnesia, happy-path burn, and latency theater, which involve treating two samplers as one agent, sharing a disk, spending free run on demos, and ignoring timeout in harness code. The source recommends spending the free probe on the ugly path first, booting a clean worktree per replay run, and recording duration as telemetry beside the result.",
  "summary": "Your free probe passed on a scratch box. That green badge still measured the wrong machine. A green free endpoint is a cheap probe. It is not a production ship gate. Treat it like a probe, or you ship luck. I keep watching agents pass on free compute. Then the same patch dies on the real model. Does that gap sound familiar to you? That gap is not a dumb model. It is an eval-transfer bug in the…",
  "key_points": [
    "Green free endpoint is cheap probe, not production ship gate",
    "Eval-transfer bug causes wrong machine measured in harness",
    "Pin frozen identity, avoid replay confusion across endpoints"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}