MobileTopUP: Designing a Reliable Recharge Transaction Workflow
Distributed transactions become interesting exactly when the happy path stops being reliable. A basic recharge implementation might look like this: charge card ↓ call recharge provider ↓ return success In a local test environment, that can appear perfectly adequate. Production introduces: duplicate requests; timeouts; delayed provider responses; asynchronous status updates; partial failures;…
When designing a reliable recharge transaction workflow, the initial implementation might seem sufficient when tested locally. However, in a production environment, various issues can arise such as duplicate requests, timeouts, delayed provider responses, and partial failures. A single user action can involve multiple components like payment, external fulfilment, asynchronous confirmation, and the operation should never be executed twice accidentally.
To address these challenges, a transaction needs an explicit lifecycle. This begins with creating a transaction record before external execution. The record contains essential details such as transaction ID, status, recipient, quote ID, idempotency key, and the timestamp of creation.
Every subsequent operation must have an internal identity, which can be used in logs, payment metadata, provider metadata, support tools, and reconciliation jobs. The use of a state machine instead of a success flag provides better clarity in handling asynchronous processing. Various states can be defined and for each state, it's crucial to understand what must have happened previously, which transitions are allowed, and whether the state is terminal. Every state should be questioned on whether a worker can safely retry from this point.
Idempotency is crucial at the transaction boundary. For instance, if a user accidentally double-clicks or the frontend retries due to a slow response, without idempotency protection, the customer might pay twice or the recipient might receive duplicate value. Instead, an idempotency key is required, which is stored with a unique database constraint.
The pseudo-code for this could be: 'existing = find_transaction(idempotency_key); if existing: return existing transaction = create_transaction(idempotency_key); process(transaction)'.
Retries should not be blindly executed on ambiguous provider requests. If a provider request is timed out, it means the outcome is uncertain, not a failure. If the provider supports idempotency, it should be used. If it provides transaction lookup, the original request can be queried. If it sends asynchronous callbacks, these need to be waited for within a reasonable reconciliation window.
Every provider operation needs a correlation identifier for traceability. For example, the internal transaction ID 'txn_123' can correspond to a provider request reference 'mobiletopup-txn_123' and the provider transaction ID 'ext_987'. This correlation helps in moving between systems without searching by amount and timestamp. Payment and recharge states should be separated. The overall user-facing status can be derived from these separate dimensions. It's important not to compress both into one column too early.
Asynchronous completion is normal and external fulfilment should work as an asynchronous workflow. An API request creates a transaction, authorizes/captures payment, enqueues a recharge job, submits to the provider, processes it, receives a webhook or polling update, and finally reaches a final state. The API does not need to maintain an open HTTP connection throughout the provider lifecycle. The frontend can fetch updates through polling or webhooks.
Webhooks too need defensive engineering. They can arrive once, twice, out of order, or much later than expected. Therefore, webhook processing should be idempotent, and the authenticity of the webhook needs to be validated using the provider's supported verification mechanism. Lastly, even if webhooks exist, reconciliation should still be performed. This ensures that events can be periodically checked for processing status using a reconciliation worker, providing an additional layer of assurance.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.