Urgent.News

What's breaking now, across thousands of outlets.

AI

One call, from fetch to a validated object — streaming structured output without the plumbing

You asked the model for a JSON object and set stream: true , because you want the UI (or the next step) to start the moment the first field lands instead of blocking on the last token. Then you're here: const res = await fetch ( url , { method : " POST " , body }); const reader = res . body ! . getReader (); const decoder = new TextDecoder (); let buf = "" , json = "" ; for (;;) { const { value ,…

You instructed the model to deliver a JSON object, enabling streaming output. You initiated the request using fetch with POST method and provided the body. A reader and decoder were set up to process the response. The loop iterated, appending decoded lines to a buffer and parsing JSON. If the payload matched [DONE], it continued to the next iteration.

After each line, the buffer was updated to exclude the processed lines. On successful parsing, the data was set or appended to the JSON string. Once the loop completed, the parsed JSON was converted into a valid object using Schema.parse. The operation entailed tackling three issues - the transport, partial parsing, and validation.

I created three separate packages to address these issues, followed by another package that integrated them. With a single call to structured, you obtained both an async iterable of partial snapshots and a holder for the final validated result. The return value is both an async iterable for live partials and a holder for the final validated result.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

No API Keys, No Cloud Bills: Running a Document Pipeline Entirely On-Device

No API Keys, No Cloud Bills: Running a Document Pipeline Entirely On-Device We run a document processing pipeline on a Mac mini. No API keys. No cloud inference.

  • Document processing pipeline runs entirely on-device without API keys or cloud inference
  • Mano-P GUI agent uses 4B parameter model with W8A8 quantization for on-device processing
  • Local 4B model achieves 56% pass rate vs 39% for cloud-based Qwen3-VL-Plus

OpenAI API Rate Limit Errors (429): Which Ones to Retry and Which Ones to Stop

Originally published on the Djangix blog: OpenAI API Rate Limit Errors (429): Which Ones to Retry and Which Ones to Stop A 429 from the OpenAI API is not one problem — it is at least two very…

  • 429 status code indicates two issues: temporary overload or quota/billing limitations
  • Prevention strategies like caching, batching, and model selection reduce long-term rate limit errors

More from Friday 9 October →