How to monitor your ISP's performance and prove SLA violations

Start 30-day free trial Try now, sign up in 30 seconds

Your servers are healthy. Your cloud infrastructure is fine. But your users are complaining, transactions are timing out, and your support queue is growing faster than your revenue. The problem isn't in your data center—it's somewhere in the invisible chain of routers, transit providers, and ISPs carrying your traffic across the internet.

This is the quiet frustration of network operations. ISP degradation rarely triggers an alarm. It creeps in as slightly elevated latency, an extra hop, or a silent route change that pushes your traffic through a congested transit network in another country. When you call your ISP provider, the answer is almost always the same: "We see no issues on our end." Without independent data, that is where the conversation ends.

It doesn't have to. With continuous ISP monitoring in place, you can track every hop your data takes, measure every millisecond of delay, and build a timestamped, evidence-backed record that transforms an SLA dispute into a provable claim.

What your ISP actually promises—and what often gets lost in translation

When you sign a contract with an ISP or transit provider, you aren't just paying for bandwidth. You're buying an SLA—a formal document that defines minimum performance standards across uptime, latency, and mean time to repair (MTTR).

The challenge is that SLAs are only as useful as the data you have to measure them against. Most ISPs measure their performance from their own vantage points, which may not reflect what your end users actually experience. Relying solely on provider-reported metrics to evaluate compliance means trusting the very party you are trying to hold accountable.

You need an independent, continuous measurement from locations that represent your actual users—capturing what is happening at every node between your origin and your destinations.

The metrics that tell the real story

Before you can prove a violation, you need to know what to measure. ISP performance isn't a single number—it's a constellation of metrics that, together, paint a complete picture of how data travels across the network.

Latency is the round-trip time of a packet from source to destination, measured in milliseconds. It's the most visible symptom of ISP degradation. When users say "the site is slow," latency is almost always part of the story.

Jitter is the variation in latency over time—the inter-packet delay variance. A connection with consistently high latency is actually more tolerable than one with unpredictable jitter, because applications can buffer for predictable delays. Jitter is the villain behind choppy video calls, stuttering VoIP audio, and erratic application behavior.

Hop count is the number of intermediate routers your data passes through between source and destination. A sudden increase in hop count often signals a route change—sometimes benign, sometimes a symptom of congestion or a peering dispute between providers.

Maximum transmission unit (MTU) is the largest packet size the network path can handle without fragmentation. MTU mismatches are a sneaky source of application failures, particularly for protocols that use large packets like certain VPN configurations.

Together, these metrics give you the full picture of ISP performance. A single snapshot isn't enough—you need continuous measurement over time, from multiple geographic vantage points, so you can distinguish a momentary blip from a sustained SLA violation.

Why multi-location monitoring is non-negotiable

One of the most common mistakes in ISP monitoring is measuring metrics from a single location. Your data center might have a perfectly good connection to your ISP's nearest point of presence, while users in another city—or another country—are experiencing a completely different path with completely different characteristics.

This is where multi-location, multi-vantage-point monitoring becomes essential. By pinging your hosts from different geographic locations simultaneously, you can isolate whether a performance issue is localized (affecting users in a specific region) or systemic (affecting everyone). You can also identify which specific path—including the combination of ISPs and transit providers—is responsible for the degradation.

When you overlay results from dozens of locations on a global latency map, patterns that would be invisible in a single-location view jump out immediately. A cluster of high-latency readings from the United States and perfectly normal readings from Europe points straight to a regional ISP or transit provider issue—not a problem with your infrastructure.

Building the evidence: Network path analysis with traceroute

If latency metrics tell you that something is wrong, traceroute data tells you where it's wrong. A traceroute maps the actual path your data takes, hop by hop, from source to destination. It reveals the IP address, organization, latency contribution, and AS number of every intermediate router along the way.

Reading a traceroute well is a skill. Here's what to look for:

  • A sudden latency spike at a particular hop, followed by consistently high latency through all subsequent hops, indicates congestion or a problem at that specific node. A hop that doesn't respond (shown as asterisks in a traditional traceroute) isn't necessarily broken—some routers are configured to drop ICMP packets silently. But repeated non-responses combined with high downstream latency can indicate a problem.
  • A change in the AS numbers along the path—compared to a baseline traceroute from a week ago—tells you that your traffic is being routed differently. This might be because your ISP made a routing policy change or a peering dispute that forced traffic onto a different transit path. Neither of these events may have been disclosed to you, but both can directly impact your users.

The ability to drill down to hop-level detail from multiple monitoring locations—and compare those paths over time—turns a traceroute from a manual troubleshooting tool into a continuous audit trail of how your traffic moves through the internet.

How Site24x7 ISP latency monitoring puts detailed analysis into practice

Site24x7's ISP Latency monitor is built specifically for this kind of continuous, multi-location network path analysis. Once configured, it polls your specified host across multiple locations at configurable frequencies—from every five minutes to every 30 minutes—and collects the full suite of metrics: availability, latency, jitter, hop count, MTU, and distinct AS number count from each location.

The Summary view presents all of these metrics in one place. A global latency map plots round-trip times from every monitoring location, giving you an immediate visual sense of where performance is strong and where it's degraded. The latency and jitter trend graphs let you scroll back through time to see exactly when a degradation started—which is exactly what you need when building a timeline for an SLA dispute.

The Latencies of all Paths section shows per-location breakdowns across all key metrics. You can see which locations report the highest hop count, which paths have the worst jitter, and which locations can't reach the host at all (tracked as Unreachable Paths).

View the individual location details at every hop with the latency details of all paths.

Clicking View Traceroute from any location pulls up the raw traceroute with latency, jitter, MTU, hop details, and IP prefix at every node.

Obtain the raw traceroute with details like latency, jitter, and MTU in bytes, number of hops, and IP prefix.

The Recent Network Path Analysis visualization takes you further into any location, rendering the traceroute from every server as an interactive graph. Nodes represent hops; edges represent the connections between them. Hovering over any connection reveals the location's agent name, latency, hop count, IP prefix, AS number, and AS name. You can highlight individual paths, set a link delay threshold to flag problematic segments, search for specific nodes, and trigger an on-demand poll from any location for a fresh real-time traceroute.

The Recent Network Path Analysis section shows the traceroute from every location server.

Site24x7's threshold-based alerting works at the hop level, covering latency, jitter, and AS numbers. Flag a specific AS as problematic and get alerted the moment your traffic routes through it or trigger an alert if the AS at the last hop changes unexpectedly. Thresholds can be set globally or per location, depending on how granular your SLA requirements are.

The monitor supports ICMP, TCP, and UDP protocols, and works with both IPv4 and IPv6 addressing. On-Premise Pollers can be used as monitoring locations for internal network paths, using Linux-based pollers specifically for ISP latency monitoring.

To learn more about the monitor and how to add it, see our help documentation.

Turning data into documentation: Proving an SLA violation

Good monitoring data must be organized clearly in order to be useful in an SLA dispute.

Start by establishing a baseline during a stable period—two to four weeks of ISP Latency monitor data that documents typical latency ranges, hop counts, and AS paths per location. Without a baseline, you can't prove a deviation.

When degradation occurs, Site24x7's alerting and reporting tools capture it automatically: timestamps, affected metrics, impacted locations, and duration. Correlate that data with traceroute captures from the same window—if the AS path changed the moment latency spiked, that linkage is your evidence.

When you present the claim, be precise. Include exact latency figures against the contracted threshold, the duration to the minute, the monitoring location it was measured from, and the AS path change. The specificity of real numbers, real timestamps, and named providers is what converts a complaint into a credible claim.

Tracking every route change

ISPs modify routing policies regularly due to peering disputes, capacity decisions, and infrastructure changes—and they're rarely required to tell you. One day your traffic takes a clean 8-hop path; the next it's making a 20-hop detour through a congested transit network. Your users feel the slowdown before anyone on your team knows the change happened.

By tracking AS number changes at each hop continuously, Site24x7's ISP Latency monitor builds a clear record of when your routing changed, where it was redirected, and what impact it had on latency and jitter, turning an invisible infrastructure shift into documented, timestamped evidence. That's the difference between reactive firefighting and having the data to hold the right provider accountable.

Watching the supply chain

No single ISP controls the entire path your data takes. Effective ISP monitoring is less like checking a single contract and more like auditing a supply chain. Every hop is a link, and every link has to perform. When one doesn't—whether it's your primary ISP, a transit provider, or a peering partner three AS numbers removed from your direct contract—your users experience the impact regardless of where the fault lies.

Site24x7's ISP latency monitoring gives you your own data, from your own vantage points, at your own frequency, with the traceroute evidence to back every claim. The next time someone on the other side of an SLA review says "We see no issues on our end," you'll have the data to steer the conversation towards a resolution that gets your site running smoothly again.

Frequently asked questions

What is ISP latency monitoring and why do I need it?

ISP latency monitoring is the continuous measurement of network performance between your infrastructure and your users tracking latency, jitter, packet loss, hop count, MTU, and AS path data at every node along the route. You need it because ISP degradation rarely triggers a conventional alert. It surfaces as slightly elevated response times, a silent route change, or an extra transit hop through a congested network—none of which your traditional server or application monitoring will catch. Independent, continuous ISP monitoring gives you the data to detect these shifts as they happen and the evidence to act on them.

What is the difference between latency and jitter, and which one matters more for application performance?

Latency is the round-trip time of a packet from source to destination. Jitter is the variation in that latency over time. For most applications, jitter is the more damaging of the two. A consistently high-latency connection is something applications can buffer around, but unpredictable jitter causes choppy video calls, stuttering VoIP audio, and erratic behavior in real-time applications. Both matter for SLA compliance, but jitter is often the metric that explains why users are complaining even when average latency looks acceptable.

What is an AS number and why does it matter for ISP monitoring?

An autonomous system (AS) is a network operated by a single organization—an ISP, transit provider, or content network—under a unified routing policy. Every hop your data takes across the internet passes through one or more ASes, each identified by a unique AS number. Tracking AS numbers at each hop matters because ISPs can silently reroute your traffic through different transit providers without notifying you. If that new transit provider is congested or poorly peered, your users experience the degradation, but without AS-level visibility, you have no way to prove where the fault lies or which provider to hold accountable.

Why is monitoring from a single location not enough for ISP performance measurement?

Your data center might have a perfectly healthy connection to your ISP's nearest point of presence while users in another city—or another country—are on a completely different network path with completely different characteristics. Single-location monitoring masks this entirely. Multi-location monitoring lets you distinguish between a localized issue affecting a specific region and a systemic failure affecting everyone, and it identifies which specific combination of ISPs and transit providers is responsible for the degradation. Without multiple vantage points, you are measuring your network's performance, not your users' experience of it.

How do I read a traceroute to identify where ISP degradation is happening?

Look for a sudden latency spike at a specific hop followed by consistently elevated latency through all subsequent hops. That hop is where congestion or failure lives. A hop that shows asterisks (no response) isn't necessarily broken, as some routers are configured to drop diagnostic packets silently, but repeated non-responses combined with high downstream latency warrant investigation. Most importantly, compare the AS path: if the AS numbers along the route have changed since your last baseline traceroute, your traffic is being through a different transit provider, and that change may not have been disclosed to you. That AS path shift, correlated with a latency spike at the same timestamp, is the core of any credible SLA violation claim.

How does Site24x7 ISP latency monitoring detect route changes I was never told about?

Site24x7 tracks the AS number at every hop continuously across all monitoring locations. When your traffic suddenly starts transiting through a different AS—one that wasn't in the path the day before—the monitor records the exact timestamp, the affected locations, and the latency and jitter impact of that change. You can configure threshold-based alerts to fire the moment a specific AS appears in your path or when the AS at the last hop changes unexpectedly. Over time, this builds a documented, timestamped record of every routing change your traffic has experienced, including the ones your ISP never mentioned.

What is the difference between a traceroute snapshot and an MTR report?

A traceroute gives you a single point-in-time map of the network path from source to destination, which is useful for identifying where a failure lives at a specific moment. A My Traceroute (MTR) report runs continuously, polling each hop repeatedly over time and building a picture of how latency and packet loss shift at each node across multiple cycles. If a specific hop consistently shows high packet loss across multiple MTR cycles, that node is a persistent fault.

How do I build a credible SLA violation claim using monitoring data?

Start by establishing a documented baseline during a stable period—two to four weeks of ISP Latency monitor data capturing typical latency ranges, hop counts, and AS paths per location. When degradation occurs, Site24x7 captures it automatically: timestamps, affected metrics, impacted locations, and duration. Correlate that data with traceroute captures from the same window—if the AS path changed at the same moment latency spiked, that linkage is your evidence. That level of specificity—real numbers, real timestamps, named providers—is what converts a complaint into a credible claim.