Urgent.News

What's breaking now, across thousands of outlets.

AI

Solving AWS re:Post's #1 GenAI Headache: Real-Time Token Streaming with Amazon Bedrock & AWS Lambda

If you have spent any time working with Amazon Bedrock over the past year, you have probably noticed a recurring conversation on AWS re:Post under the #AmazonBedrock and #AWSLambda tags. Every couple of days, an engineer asks the exact same question: "I hooked up Bedrock (Claude 3.5 Sonnet) to a Lambda function behind API Gateway, but my chat app takes 10 to 15 seconds before the first word shows…

If you have worked with Amazon Bedrock over the past year, you may have noticed a recurring question on AWS re:Post: How do I stream tokens as they are generated when integrating Bedrock with Lambda behind API Gateway? The underlying issue is that API Gateway buffers the entire response in memory, causing significant delays in displaying the chat app's first word.

Additionally, API Gateway has a 29-second hard limit for integration timeouts, which can terminate the connection if the model takes longer than expected to generate a response. This often leads to a frustrating user experience with perceived latency of 8 to 15 seconds before any words appear. The architecture typically involves a web browser sending an HTTP request to Amazon API Gateway, which then forwards the request to an AWS Lambda function.

The Lambda function interacts with Amazon Bedrock to generate responses, but the full buffering and timeout issues break the token streaming capability that Bedrock offers out of the box. The solution is to use Lambda Function URLs in RESPONSE_STREAM mode, which allows Lambda to send data back to the client using HTTP chunked transfer encoding.

This means that bytes leave Lambda and hit the browser immediately as they are generated by Bedrock, providing real-time token streaming with sub-300ms Time-To-First-Token (TTFT).

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

More from Saturday 3 October →