One pass of my eval bills $9.14 on the API and $0 through the CLI
One pass of my board eval bills $9.14 on the Anthropic API. Through Claude Code it bills $0. Same model, claude-opus-4-8. That is 27 calls, and it is not an estimate. The CLI prints a total_cost_usd in its envelope: what the run would have cost on the API. It bills the subscription instead, so the number is a receipt for money nobody spent. The switch fixed something better than the bill Running…
One pass of the eval suite costs $9.14 when executed via the Anthropic API, as opposed to $0 when run through the Command Line Interface (CLI). With 27 calls made, the evaluation was not estimated. The CLI reveals the total cost through the envelope, displaying the expense that would have been charged to the API during a run. By utilizing the subscription instead, the bill represents money never spent.
The CLI switches to a structured output format identical to that of the API, using an internal tool call. Initially, format reliability was at 7 out of 15, but improved to a perfect 15 out of 15. However, the CLI validator rejects certain draft elements, including pattern , minLength , maxLength , minItems , maxItems , format, and the $schema meta-ref.
Nonetheless, the strict Zod validation in the caller side remains intact, with validation shifting to a location where it can function effectively. A critical issue arises when ANTHROPIC_API_KEY is present in the child process environment; the CLI inadvertently bills the API account rather than the subscription. This silent money leak remains unmarked by errors or warnings, with the invoice arriving at the end of the month for a run believed to be free.
As a developer loop tool, Anthropic's consumer terms restrict automated access, except when utilizing an Anthropic API Key or when explicitly permitted. Commercial terms governing API use do not extend to consumer subscriptions. The CLI is expected to be used as designed within this context. A commercial service, however, would not be covered under these terms.
It is crucial to understand the true cost of evaluating one's own code and to be aware of any potential money leaks.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.