ICYMI: What landed for AI builders in August 2026
A recap of August 2026 launches for AI builders across Amazon Bedrock, Amazon Bedrock AgentCore, and Strands: million-token context for OpenAI models, cross-Region inference, agents that run for up to 14 days on dedicated compute, expanded AWS GovCloud availability, and Strands Robots for physical deployment.
Amazon's latest updates to its AI offerings in August 2026 aim to provide builders with increased capabilities and control. Amazon Bedrock, already used by over 225,000 customers, now includes AgentCore, enabling users to build, connect, and optimize agents using any framework and model. The Strands Agent Harness SDK has been released as open source, further enhancing flexibility for creating and deploying agents.
Key enhancements include:
- Million-token context windows for OpenAI applications on Amazon Bedrock, with GPT-5.6 Sol, Terra, and Luna supporting prompt caching to reduce cost and latency.
- Web Search functionality allowing models to find current information beyond their training data, incorporate results into answers, and provide citations without separate search integrations.
- Cross-Region inference, enabling access to GPT-5.6 Sol, Terra, and Luna from over 25 AWS Regions, with global profiles offering broader capacity at lower per-token prices.
- IAM principal cost allocation and AWS Cost Anomaly Detection to attribute inference spend, monitor third-party foundation model costs, and provide root-cause breakdowns for unexpected changes.
- Faster cyber defense through Daybreak Red and Blue from OpenAI, offering defensive and authorized offensive workflows with advanced identity verification, monitoring, access controls, and zero-operator-access infrastructure.
- Extended AgentCore runtime instances on dedicated Amazon EC2 compute, with sessions lasting up to 14 days, including GPU-accelerated, memory-optimized, and compute-optimized instances.
- Temporal policies for evaluating agent actions against previous behavior, enabling enforcement of sequences, prerequisites, approval gates, matching values between calls, and data freshness.
- Rate limiting for controlling requests, inference tokens, and concurrent connections by user or group.
Written by urgent.news from AWS Machine Learning's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.