Measuring the Rust Storage Engine: Lioran S3 PUT Timing from Network to RocksDB
Measuring the Lioran S3 Rust Storage Engine A benchmark that says: WRITE: 400 MB/s creates more questions than it answers. Where did time go? Network? SHA-256? Filesystem writes? fsync ? Directory creation? Rename? RocksDB metadata? Lioran S3's normal PUT path contains explicit timing instrumentation for these stages. I’m Swaraj Puppalwar , Founder & CTO of Lioran Group / Lioran Developer…
Measuring the performance of the Lioran S3 Rust Storage Engine reveals various factors that contribute to its storage operations. The benchmark highlights WRITE operations at 400 MB/s, but this raises more questions than answers. Several components are examined, including network, SHA-256, filesystem writes, fsync, directory creation, rename, and RocksDB metadata.
The Lioran S3 engine provides explicit timing instrumentation for these stages, enabling developers to trace the various timings involved. The engine checks for acceptance of trace timings through the BASTION_TRACE_PUT_TIMINGS setting. Streaming timings measure the duration of various stages, such as recv_duration, write_duration, sha256_duration, flush_duration, fsync_duration, total_stream_duration, close_duration, mkdir_duration, rename_duration, and metadata_duration, leading to a total_duration for the PUT operation.
A conceptual decomposition of a PUT operation can be represented as follows: TOTAL PUT ├── receive from AsyncRead ├── SHA-256 ├── staging writes ├── flush ├── fsync ├── close ├── mkdir ├── rename └── RocksDB metadata commit.
This breakdown helps in identifying the dominant stage during the PUT operation for performance tuning purposes. There are four cases to consider:
1. Receive dominates the operation. If recv_ms is 900 ms, write_ms is 100 ms, sha256_ms is 30 ms, and fsync_ms is 20 ms, the server is mostly waiting for incoming bytes. In this case, optimizing RocksDB won't fix the issue. Instead, focus on factors like client upload speed, network, TLS, reverse proxy, buffering, and request-body behavior.
2. Fsync dominates the operation. If recv_ms is 100 ms, write_ms is 80 ms, and fsync_ms is 700 ms, the issue lies in durability synchronization. Questions include storage device latency, filesystem virtualization, durability mode, and write cache behavior.
3. Metadata dominates the operation. If metadata_ms grows significantly under load, investigate the metadata path within RocksDB. This includes examining RocksDB compaction, WAL behavior, block cache, write stalls, disk contention, and concurrent metadata operations.
4. Hashing (SHA-256) matters. SHA-256 is calculated incrementally on every incoming chunk. At very high throughput, the CPU cost of checksumming can become measurable. However, removing checksums is not the solution; instead, measure the actual cost before concluding there is a bottleneck.
When publishing benchmark results for Lioran S3, it is essential to record information such as CPU, RAM, disk, filesystem, operating system, container/native environment, network topology, reverse proxy, TLS settings, object size distribution, concurrency levels, and software version. This ensures that benchmark runs are measuring the same system and allows for proper optimization.
A sane optimization loop involves measuring performance first, identifying the dominant stage, making changes to address it, and then measuring again. This iterative process ensures that performance work remains evidence-driven and prevents unnecessary optimization efforts.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.