{
  "id": 11995557,
  "title": "Testing the tool call after the HTTP success",
  "url": "https://urgent.news/2026/10/04/testing-the-tool-call-after-the-http-success",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-04T20:07:02.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/arthur031221/testing-the-tool-call-after-the-http-success-4ma"
  },
  "original_language": "en",
  "account": "The source material describes a Python command-line tool called toolcall-check that validates the behavior of tool calls made through a chat completions compatible endpoint. Even though HTTP 200 status code indicates success, it does not guarantee that the client received the tool call correctly. The tool performs five different probes to test various scenarios where the tool call might be problematic. These scenarios include forced echo calls, nested JSON calls, streaming versions of both calls, and a two-turn echo round trip check. The tool also checks for specific types of data such as integers, booleans, decimal numbers, lists, and exact object keys within the nested payload. It verifies that the value and JSON type match, treating true, 1, and 1.0 as distinct. For streamed responses, the tool parser can handle fragmented UTF-8, CRLF, comments, multiline data fields, and tool call deltas keyed by choice and call index. The tool requires a specific \"finish reason\" and a \"[DONE]\" marker in the streamed responses. It also checks for duplicate JSON keys and nonfinite values, which would cause the validation to fail. The tool's HTTP worker operates in a separate subprocess, with a parent wall time limit to handle slow DNS, headers, and trickle responses. Redirects are not followed, response bodies are limited in size, and authorization headers are not included in the trace. The report generated by the tool is a self-contained HTML file with escaped evidence and no scripts or external requests. It redacts secrets as much as possible, advising users to inspect the report before sharing it. The tool includes a synthetic demo that exercises the real HTTP client and writes artifacts similar to what would be produced by the endpoint itself, but it does not confirm that a remote model or serving stack meets these requirements. The project focuses on the integrity of nested arguments, streamed fragment reconstruction, the fixed result in the second turn, and the inspectability of traces. The author has not yet validated these probes against remote models. The tool has no runtime dependencies and generates three files: report.html, results.json, and trace.json, placed in a new private directory. The author is seeking feedback on the strict stream markers, response shapes that should be supported, and potential local endpoints to check.",
  "summary": "HTTP 200 doesn't mean the client got a good tool call. A nested integer can arrive as a string, a streamed call can lose a fragment, or a later assistant turn can ignore the tool result. The traces for these protocol mistakes can be awkward. I built toolcall-check, a Python CLI that checks these behaviors on a chat completions compatible endpoint. It runs five probes: a forced echo call, a forced…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}