{
  "id": 11713653,
  "title": "Day 5: We scanned a 200-route codebase into OpenAPI without uploading a file",
  "url": "https://urgent.news/2026/10/03/day-5-we-scanned-a-200-route-codebase-into-openapi-without-uploading",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-10-03T15:42:52.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/jeff_pdc/day-5-we-scanned-a-200-route-codebase-into-openapi-without-uploading-a-file-1k2l"
  },
  "original_language": "en",
  "account": "Days 2 through 4 of the series focused on scanning a new API with a spec first. Most teams operate in the opposite scenario, where the service is three years old, the original documentation is outdated, and the existing code contains over 200 routes that still support production. On Day 5, the series explores the process of scanning existing code into an OpenAPI document that is accurate enough for trust. Recording traffic through a proxy is fast but provides an incomplete picture, as it only captures the paths that are actually used. Error branches, admin routes, seasonal endpoints, and rarely triggered endpoints like 409 errors remain invisible to the proxy. Writing the specification by hand yields the best document, but it only lasts a week as the code evolves rapidly. Pointing AI at the code repository works well for smaller demonstrations but struggles with larger, production-scale codebases. A dedicated AST (Abstract Syntax Tree) scanner, combined with human review, proved to be the most effective approach. Parsing code through a syntax-tree walk reveals the intricacies that a simple text search often misses, especially when routing is nested across multiple routers or handled by decorator-based frameworks like Spring's @RequestMapping, FastAPI's path operations, or ASP.NET's attribute routing. Schemas can be derived from type systems when available, such as TypeScript interfaces, Pydantic models, Java DTOs, Go structs, and C# records. This method produces complete operations deterministically, without the need for AI assistance. The scanner, developed by Powerduck, supports eight programming languages including TypeScript, JavaScript, Python, Go, Java, C#, Rust, PHP, and more. It covers various frameworks such as Express, Fastify, NestJS, Koa, Hono and Next.js, FastAPI, Flask, Django REST, Starlette, Gin, Chi, Echo, Fiber, net/http, Gorilla mux, Spring, JAX-RS, Micronaut, ASP.NET, FastEndpoints, Axum, Actix, Rocket, Laravel, Symfony, and Slim. Both HTTP and Server-Sent Events (SSE) are supported as first-class outputs. Honest gaps are preferred over confident guesses, as they ensure the output remains trustworthy. When the scanner encounters unknown elements, it flags them as gaps instead of fabricating plausible schemas. For each gap, an optional AI resolver can generate a proposed schema for human approval. By default, the scan operates without AI assistance, maintaining a purely human-driven approach. Turning on the AI step increases recall for loosely typed code but keeps the human review gate in place. The safety of proprietary code is maintained through local scanning, with no source code uploaded to external servers. The optional AI call is also opt-in, applied individually to each gap and scoped to a single handler. The initial scan captured several important features, such as a controller returning different DTOs based on status codes, resulting in two explicit response schemas. Another example involved an SSE endpoint documented with a generic text/event-stream schema, which was later replaced by a more specific schema. Additionally, a route registered in two places with different middleware was flagged to highlight an authentic authentication inconsistency. Several query parameters were identified that did not exist in the previous Postman collection, indicating routes that were never explored. Rescans are performed as differences, allowing the scanner to adapt over time. After the first successful import, the scanner generates a discovery sidecar file within the workspace. Subsequent rescans produce change sets, detailing new routes, removed endpoints, and renamed parameters. Existing descriptions, examples, and manual edits from the initial pass are preserved, ensuring a continuous evolution of the OpenAPI document. This approach transforms the initial documentation process from a time-consuming project to a regular rescan, review, and merge procedure. The scanner clearly communicates its limitations to users, helping them understand the scope of coverage. It reports routes that cannot be scanned, providing method details, reconstructed paths, confidence levels, and lists of gaps. For gaps that require filling, an optional AI resolver can generate a proposed schema, but the final decision remains in the hands of human reviewers. The scan itself runs entirely on the local machine, ensuring that source code remains private. AI assistance, when enabled, is also an opt-in feature, scoped to individual handlers, providing an additional layer of control. Overall, this scanning process offers a reliable and efficient way to maintain accurate OpenAPI documentation, adapting to the ever-changing landscape of software development.",
  "summary": "Days 2 through 4 covered the easy case: a new API where the spec can come first. Most of us live in the other case. The service is three years old, the original team is gone, the docs are a Postman collection from 2023, and two hundred routes keep production running. Day 5 of the the series is about reversing the direction — scanning existing code into an OpenAPI document that is accurate enough…",
  "key_points": [
    "Powerduck scanner scans 200-route codebase into OpenAPI without uploading files",
    "AST scanner combined with human review proves most effective for large codebases",
    "Scanner maintains privacy by running entirely on local machine"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}