Urgent.News

One page, thousands of outlets. See who else covered it.

Editions

AI

Enterprises are overpaying for simple AI queries — Snowflake's gateway now auto-routes to cut costs up to 3x

Enterprise teams running AI agents at scale are finding that a single model handles every task poorly — either the model is too expensive for simple questions or not capable enough for hard ones. Model routing, which picks the right model for each task automatically, is becoming the fix. Snowflake’s Cortex AI Gateway now offers dynamic model routing to address that: enterprises can select “auto”…

Enterprises are overpaying for simple AI queries — Snowflake's gateway now auto-routes to cut costs up to 3x

Enterprise teams that run AI agents at scale are discovering that a single model struggles with various tasks. This is because the model may be too costly for simple questions or not powerful enough for complex ones. Snowflake, a cloud data and analytics company, has introduced a solution called the Cortex AI Gateway, which employs dynamic model routing to tackle this issue.

With this feature, enterprises can opt for the "auto" setting instead of choosing a specific model, allowing the system to automatically route each task to the most suitable model based on a combination of quality and cost. According to Snowflake, this functionality can potentially reduce token costs by up to three times for certain workloads, as the company's internal tests revealed that simple queries were often handled by its most capable model, resulting in unnecessary expenses.

The company emphasizes that model routing is more than just a price and performance consideration; it also involves governance and context. Baris Gultekin, Snowflake's vice president of AI, stated that for high-quality, enterprise-grade AI agents to be built, context and governance must be addressed. The new capability builds upon Snowflake's Cortex AI Gateway, which was launched in July 2026 as a governance layer for agent and model traffic.

The dynamic routing process consists of two mechanisms: a small model first attempts the task, known as the advisor pattern. If the smaller model fails to complete the task, it calls upon a larger model to assist. A classifier sorts the tasks based on historical data and automatically assigns straightforward questions to simpler models.

Customers can still designate a specific model if desired, but auto-routing is optional. Pricing for AI remains straightforward, being solely based on token usage. If routing to a more cost-effective model occurs, the bill will be lower, without any additional charge for the routing decision itself. Access controls adhere to the task rather than just the data.

Snowflake ties routing to the same access controls used for data governance, starting at the data level with role-based access controls. These controls then extend to models, where customer roles are mapped to approved model buckets, and finally to agents, where privileges can be restricted to narrower permissions than the user invoking the agent.

Open models can be run from the customer's region to comply with data residency requirements. Snowflake's integration of Natoma adds an additional layer, bringing more than 100 MCP connectors with scoped, governed access. Agents can be granted read-only access to connected tools like email, rather than broader permissions. Context plays a crucial role in enabling cheaper models to handle tasks effectively.

Snowflake recently unveiled its Horizon Context and Cortex Sense tools, which provide context capabilities. Without proper context, a model must perform exploratory work, such as writing and testing SQL, searching data, and retrying when necessary, which is expensive and typically requires a more capable model. By packaging context in advance, simpler, cheaper models can often handle the same tasks.

Snowflake also incorporates agent memory into the context, allowing the system to remember previous interactions and avoid re-solving the same problem from scratch each time.

Written by urgent.news from VentureBeat's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Also reported by 1 other outlet

Read the original at venturebeat.com →

More in AI

More from Tuesday 18 August →