Uber Eats Rebuilds Search Pipeline to Cut End-to-End Latency by 50%
Uber has rebuilt major parts of the Uber Eats search pipeline, reporting a 50% reduction in end-to-end latency. Changes include Above-the-Fold measurement, reduced retrieval work, parallel hydration, advertising data redesign, infrastructure optimizations, and an agentic coding workflow. Uber is also exploring microbatching, product-based retrieval, and HTTP multipart streaming. By Leela Kumili
Uber has undertaken a major overhaul of the Uber Eats search pipeline, resulting in a 50% reduction in end-to-end search latency. This initiative encompasses various aspects of the pipeline, including retrieval, feature hydration, ranking, advertising, presentation, and infrastructure. The team employed an agentic coding workflow to identify, benchmark, and validate additional optimizations throughout the process.
Initially, the focus shifted from measuring backend API response time to Above-the-Fold completion, which quantifies the time taken for the first screen of results to render with images. This shift led to significant improvements, with the changes yielding a reduction of over 200 milliseconds in above-the-fold latency.
To achieve this, Uber implemented several key optimizations. By reducing retrieval work, tens of thousands of candidates were hydrated before ranking, and many were subsequently discarded. Removing low-value retrieval strategies managed to cut about 120 milliseconds, while product-level embeddings substantially reduced data lookups by over 100 times, saving an additional 50 milliseconds.
Separating ranking hydration from presentation data resulted in a latency reduction of more than 100 milliseconds, with further optimizations such as dependency removal and request hedging contributing an extra 35 and 40 milliseconds, respectively. The advertising path was also redesigned, incorporating column-oriented bid data, in-memory access, and reduced serialization, which trimmed about 130 milliseconds from the overall latency.
In addition to these enhancements, infrastructure changes like parallel encoding, smaller embeddings, improved connection management, and Go data structure modifications aimed at minimizing garbage collection overhead, were also implemented.
The team's performance optimization approach resonated with engineers who viewed it as a model based on the Measure, Identify, Fix, Validate loop. Performance experts emphasized that the result stemmed from incremental optimizations rather than a single architectural change, highlighting the comprehensive nature of the optimizations across the entire stack.
Uber's performance experts distilled these changes into three core principles: doing less work, starting work earlier, and removing unnecessary dependencies. These principles were further connected to Uber's planned microbatching approach, drawing parallels to techniques used in AI systems to reduce synchronization between processing stages.
These changes build upon Uber's existing search platform, which has previously relied on Apache Lucene, Spark-based indexing, Kafka-based streaming updates, and a distributed serving layer. The company is now exploring end-to-end microbatching, product-based retrieval, Zero Pass Ranking, and HTTP multipart streaming. Early testing of the product-based search approach has reportedly yielded more than a 50% reduction in p99 latency.
Written by urgent.news from InfoQ's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.