A bandwidth alert is only useful if it fires at the right moment for the right reason, and getting there takes more than picking a percentage and applying it everywhere. Because traffic patterns differ wildly between a core switch, a branch WAN link, and a wireless controller, the same static threshold that works for one will either miss real problems on another or bury the team in noise from normal daily peaks. Setting up alerts that a team can actually trust means working through a handful of decisions in order, rather than configuring one rule and hoping it holds up everywhere.
1. Start with what you're actually trying to catch
Before choosing a number, it helps to decide what kind of problem the alert is meant to surface, since that answer changes everything downstream. An alert meant to catch outright link saturation needs a different design than one meant to flag a single host or application quietly consuming a disproportionate share of bandwidth. Interface-level alerts flag congestion on a specific link, while application- or host-level alerts flag the underlying traffic pattern responsible for it, for instance a single backup job, a misconfigured device stuck in a retry loop, or one user's workstation pulling far more than its share, so the team knows not just that a link is under pressure but what's actually driving it. Most environments end up needing both rather than choosing one over the other.
2. Establish a baseline before you set a threshold
A threshold set without reference to how a link actually behaves is really just a guess, and guesses tend to be wrong in one of two directions. Set it too low, and every predictable peak, a Monday morning backup job or an end-of-month reporting run, triggers an alert that the team quickly learns to ignore. Set it too high, and a genuine problem can go unnoticed until users are already affected. Watching a link's traffic over a representative period, roughly two to four weeks is usually enough to capture a normal week alongside at least one known peak event like a month-end close, gives you the baseline that makes the eventual threshold meaningful rather than arbitrary.
3. Choose the metric you're actually alerting on
A threshold is only as useful as the metric it's applied to, and "bandwidth" isn't a single number. Interface utilization (the percentage of a link's total capacity in use) is the most common starting point, since it's simple and maps directly to "is this link full." Inbound and outbound traffic, tracked separately, matters wherever a link's upload and download patterns differ meaningfully, a branch office pushing backups out overnight looks very different from one pulling large downloads during the day. Throughput, the actual rate of data moving in real time, catches problems that a longer-interval utilization average can smooth over and miss. Traffic volume, the total data moved over a period rather than a rate, is better suited to capacity planning and billing-style thresholds than to real-time alerting. Application consumption and host consumption narrow the focus further, flagging when a specific application or device crosses its own expected share rather than waiting for the link as a whole to become the problem. Picking the metric that actually matches what you're trying to catch, before setting a number against it, is what keeps the rest of this process meaningful.
4. Decide whether a spike or a pattern is what matters
Not every link needs the same kind of trigger. On a core link supporting business-critical traffic, even a brief, momentary spike might be worth flagging immediately, since a short-lived surge on the wrong link can still disrupt latency-sensitive applications like VoIP, where even a few seconds of congestion shows up as choppy audio or dropped calls rather than something users simply wait out. That calls for an instantaneous threshold, one that fires the moment utilization crosses the line, rather than waiting to see if it holds. On a link with naturally bursty traffic, though, reacting to every brief spike produces constant noise, and what actually matters there is traffic staying elevated for a sustained period rather than any single peak, which calls for the opposite configuration: a threshold that only fires once utilization has stayed above the line for a set duration, filtering out the momentary bursts that don't represent a real problem. Matching the trigger type, instantaneous or sustained, to the link's own behavior, rather than applying one rule everywhere, is what keeps the resulting alerts meaningful.
5. Build in severity, so not everything looks equally urgent
Once the trigger conditions are set, the next decision is how loudly each one should announce itself. A link edging toward its threshold during a known peak window isn't the same emergency as a WAN link that's flatlined, and treating both the same way trains a team to treat every alert as equally dismissible. Assigning severity tiers, so a minor early warning looks and feels different from a critical breach, keeps the team's attention calibrated to what the situation actually demands.
6. Route each severity to where it will actually get seen
An alert that fires correctly but lands in a channel nobody checks is functionally the same as no alert at all. Lower-severity warnings might reasonably sit in a dashboard or a daily digest, while anything approaching a critical breach probably needs to reach a person directly, whether through email, SMS, or a messaging platform the team already watches. For environments running a formal incident process, feeding high-severity alerts into an ITSM platform so a ticket opens automatically closes the gap between detection and response, rather than depending on someone happening to notice a graph in time.
7. Revisit the thresholds as the network changes
None of this is a one-time setup. As new applications get adopted, as traffic shifts toward cloud or SaaS destinations, and as the network itself grows, the baselines that were accurate six months ago may quietly stop being accurate at all. Periodically reviewing thresholds against current traffic, rather than assuming the original configuration still fits, is what keeps alerts useful instead of letting them drift into either constant noise or dangerous silence.
Setting this up with NetFlow Analyzer
NetFlow Analyzer supports this whole process directly rather than forcing a single static rule onto every interface. Thresholds can be configured per link based on that link's own traffic history, severity tiers determine how an alert is surfaced, and breaches can route straight into email, SMS, Slack, or an ITSM platform, with ServiceNow ticket creation available for anything that needs a formal response. This allows teams to configure alerting around the traffic characteristics and operational importance of each link, rather than applying one threshold universally. Get NetFlow Analyzer now to configure bandwidth alerts around how your own network behaves.
