Failure Windows in Lioran S3 V1: What the Current Rust Commit Pipeline Guarantees
Failure Windows in Lioran S3 V1 Storage-engine design becomes interesting at the word: crash. Happy-path code is only half the system. The current Lioran S3 V1 Pre-Alpha write path has a clear sequence of staging, promotion, metadata commit, and rollback. That sequence also defines its failure windows. I’m Swaraj Puppalwar , Founder & CTO of Lioran Group and Lioran Developer Solutions . This…
The Lioran S3 V1 storage engine's failure windows are important to consider due to its crash-consistency model. The current pre-alpha write path involves a sequence of stages: bucket/key validation, capacity/quota check, staging file creation, payload streaming with SHA-256, fsync in strict mode, closing the staging file, re-checking capacity/quota, creating a final parent directory, renaming the staging file to the final object path, building ObjectMetadata, writing metadata to RocksDB, and handling metadata write failures.
The most crucial aspect is promoting the physical file before committing metadata. Failure during streaming can result in removal of the staging file and no committed metadata, while rename failures lead to removal of the physical file and removal of staging cleanup. However, a crash between the rename and metadata commit can leave physical bytes without corresponding metadata records, causing orphaned disk usage.
This problem arises because filesystem namespace and embedded metadata database are not a single shared transaction. Documenting the real order of operations, rather than the intended invariant, is more valuable in the current pre-alpha phase. Future improvements could include concepts like durable publish intent recovery, journal startup reconciliation, orphan scanning, transaction state records, and idempotent repair.
However, these are design directions, not guarantees of the current V1 implementation. Pre-alpha is the right time to discuss these failure states, as it allows for kill-9, power-loss simulations, disk-full testing, metadata failure injection, rename failure injection, orphan detection, checksum verification, and restart loops to be performed.
The goal is to make every failure state defined and recoverable, rather than claiming the system is atomic and never fails.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.