We used a database as a message queue. Now we use Kafka
FoundationDB is the only database utilized by the company, a surprising fact considering its limited features as a simple key-value store. The database manages various aspects, including tenants, object metadata, and data replication across regions. Additionally, FoundationDB is repurposed as a queue to handle asynchronous tasks, much like Apple's QuiCK system used for CloudKit.
This setup has proven effective, as it avoids the dual-write problem that arises when keeping the queue within the database, ensuring all transactions remain within the database and eliminating the need for custom code to reduce read and write load. However, there are challenges associated with using a database as a message queue.
The company has experienced issues, such as the need to move asynchronous tasks like garbage collection to Kafka to alleviate the read and write load on FoundationDB. Furthermore, managing the dual-write problem becomes complicated when multiple producers are involved, as it may lead to conflicting IDs for queue items, resulting in potential data loss.
Apple has addressed this issue by implementing a robust queuing system called QuiCK, which leverages FoundationDB’s Record Layer to create a message queue based on time. This innovative solution attaches an observed time to each job, ensuring a consistent ordering of events, even in a distributed system where each component has its own independent observation of time.
The message queue lifecycle includes identifying jobs by their task type, item space, execution time, priority, and a unique ID. This allows the system to handle worker failures, crashes, and data loss without permanently losing queued items, preventing issues such as missing objects for users accessing different regions.
Written by urgent.news from Lobsters's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.