*"How I Fixed a 30-Second Timeout in a High-Latency API Call (Without Throwing Away the Data)"*
How I Fixed a 30-Second Timeout in a High-Latency API Call (Without Throwing Away the Data) The Problem: Unpredictable API Latency A client’s ASP.NET Core API was failing intermittently on a third-party payment processor with 30-second timeouts . The latency was inconsistent—sometimes 50ms, sometimes 30 seconds. The business couldn’t afford to lose transactions, but blind retries risked duplicate…
The issue at hand involved a client's ASP.NET Core application intermittently encountering 30-second timeouts when interacting with a payment processor API. The inconsistency in latency, ranging from 50 milliseconds to a full 30 seconds, posed a significant problem for the business as it prevented transactions from being completed. Retrying blindly was not a viable option due to the risk of generating duplicate payments, which would be highly detrimental to the business.
The root cause of the problem was the unreliable performance of the third-party payment processor. Attempts to automatically retry failed transactions were initially ruled out because the system couldn't afford duplicate payments. Additionally, there was no mechanism in place to recover any data that had been lost due to failed transactions.
To address this challenge, the solution implemented a time-based retry mechanism in conjunction with idempotency keys. The Polly retry policy was configured to wait for fixed intervals of one second between retry attempts. This approach was chosen over exponential backoff since the latency was unpredictable and thus, a fixed delay method was more suitable. Each retry attempt was logged to provide visibility into the system's behavior.
To prevent duplicate payments, the solution incorporated idempotency keys. Each API request included a unique identifier that was sent in the request headers. The payment processor was set up to recognize and disregard any duplicate requests that shared the same idempotency key. This ensured that even if retries were triggered, the same payment would not be processed twice.
For handling transactions that could not be automatically retried, such as those resulting from permanent errors, a Dead-Letter Queue (DLQ) was established. This provided a way to store failed transactions for later investigation and manual intervention. The system logged failed requests to a DLQ, such as an Azure Storage Queue or a database table, allowing for a record to be kept of these failures for future review or reprocessing.
Additionally, the usage of a CancellationToken was introduced to gracefully handle timeouts. By setting a cancellation token with a buffer time of 25 seconds, tasks that did not complete within this timeframe would be canceled, preventing them from hanging indefinitely. The response from the API was then awaited using this cancellation token.
The effectiveness of this strategy was measured by an 80% reduction in timeout occurrences. This improvement was attributed to the effective handling of transient failures through retries while ensuring that the system remained safe from generating duplicate payment records. Failed requests were no longer silently discarded but were instead funneled into a DLQ for manual review, thus ensuring that no transaction was lost without being addressed.
The key takeaways from this implementation emphasized the importance of choosing between retries and dead-letter queues based on the nature of the failure. Retries were deemed appropriate for transient issues where the operation is idempotent and delays are tolerable. In contrast, failures that are permanent or non-idempotent should be funneled into a DLQ for further review, ensuring that the system can recover from failures without compromising data integrity or business logic.
In practical terms, when developing code that interacts with external APIs, it is essential to always include CancellationToken to prevent tasks from hanging. Idempotency should be prioritized over retries whenever there is a risk of generating duplicate entries, especially when such duplicates could have severe business implications.
Logging failures to a DLQ is a robust strategy for ensuring that no transaction is lost and that all failures are accounted for and can be reviewed later. Regular monitoring and adjustment of retry policies are also crucial to adapt to the changing performance characteristics of the external services being interacted with.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.