Urgent.News

What's breaking now, across thousands of outlets.

Tech

Your VDS Is Not “Dropping Packets” Until You Can Prove Where

A practical workflow for diagnosing Linux VDS network problems using MTR, ss, ethtool, tc and iperf3 before escalating to your provider.

Your VDS Is Not “Dropping Packets” Until You Can Prove Where

To avoid misdiagnosing network issues, it is essential to avoid jumping to conclusions such as declaring that a provider is dropping packets without proper evidence. The source material emphasizes that an intermittent network problem could originate from various places within a system, including the guest network stack, virtual NIC, hypervisor, host interface, edge router, or upstream networks. Additionally, the issue could even stem from the application itself while the network is functioning correctly.

To properly investigate and isolate the root cause, the source recommends defining the incident clearly by noting down relevant details like source network, destination IP, protocol, port, UTC start and end times, and user-visible effects. This helps create a strong foundation for a focused investigation. It is also crucial to capture the starting state of the system by creating a private directory with a timestamp, recording various system information, deployment details, firewall changes, backups, traffic spikes, and kernel updates.

This ensures a reproducible starting point for the investigation and helps correlate events with a common clock.

The investigation process starts by running mtr tests, which help determine whether loss reaches the destination. This provides a low-impact report from the VDS to a controllable destination. It is essential not to accuse an intermediate router solely based on its high Loss% column, as routers often give control-plane replies lower priority than forwarded traffic. Different TTLs use different probes and may trace separate packet paths, so traceroute alone cannot pinpoint the faulty link.

Next, the source suggests using ss to separate the application from the path during the incident. By inspecting socket state, one can identify if an application is not consuming data promptly or if data is not being acknowledged quickly enough. This helps narrow down the layer where the problem lies, whether it is the application, guest, host, or upstream networks.

Monitoring CPU pressure alongside socket information can also provide insights into whether the guest is experiencing CPU pressure, which may delay packet processing and application work without generating clean "network error" counters.

Measuring interface counter deltas is another crucial step in the investigation. By comparing the active interface's statistics before and after a symptom or controlled test, one can identify any changes that might indicate a problem. However, it is important to note that counter meanings depend on the device and driver, so the source recommends recording ethtool settings, TSO, GSO, and GRO configurations, and avoiding disabling them as speculative fixes without proper testing.

Additionally, it is crucial to perform reverse traceroute tests from remote probes and repeat the investigation from other access networks, especially if complaints are regional, to account for asymmetric routing.

In summary, the source material advises a systematic and evidence-based approach to network investigations, avoiding premature conclusions and focusing on identifying the root cause across different layers of the system. This ensures a more accurate and effective resolution to network issues.

Written by urgent.news from HackerNoon's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at hackernoon.com →

More in Tech

More from Monday 5 October →