20× the CI traffic without getting slower: How we rebuilt Git serving at Datadog
Git fetch operations consume significant server CPU resources, especially when handling large monorepos with hundreds of thousands of files. At Datadog, CI fetches code millions of times weekly across thousands of repositories, with the largest monorepos causing fetches to take several seconds on server CPUs. This led to frequent CI outages and increased CI runtimes.
Previous attempts to scale the backend or add nodes did not reduce per-node CPU usage, as adding nodes only increased replication overhead. The solution was to adapt the Git serving architecture to handle the high volume of fetch requests efficiently. Git’s object database consists of immutable objects like trees, commits, and references, stored either as loose objects or grouped in packfiles.
Packfiles are immutable, allowing safe concurrent reads. A Git fetch operation involves negotiating which objects to send, constructing a packfile with those objects, and sending it to the client. The process can be CPU-intensive and time-consuming, especially when serving fetches for large, dense monorepos. The gitretriever solution avoids the costly packfile construction by pre-populating a Git mirror with the necessary objects, serving fetches directly from the mirror.
This approach resulted in a 20x increase in traffic without a significant increase in latency or CPU usage, demonstrating the effectiveness of the new architecture in handling Datadog's massive CI load.
Written by urgent.news from Lobsters's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.