# Network traffic alerts and thresholds: Best practices An unmonitored network fails silently, and a badly alerted network fails loudly while everyone has learned to ignore it. The distance between those outcomes is alert design: which conditions page, at what thresholds, with what severity, reviewed on what schedule. This page collects the practices that keep traffic alerting trustworthy, drawn from how experienced teams actually run it. ## In this guide: - Static thresholds, baseline deviation, and where each belongs. - Severity tiers that map to response, and the alert fatigue math. - The review loop that keeps an alerting system honest over time. ## The purpose test Every alert exists to trigger a response. That single sentence sorts most alerting decisions: if no one would act differently upon receiving it, it is a report, and reports belong on schedules rather than in pagers. Teams drowning in alerts almost always turn out to be paging their reports. ## Static thresholds: Where they belong Static thresholds fire at fixed values, and they belong wherever fixed truths exist. Link capacity is a fixed truth: sustained utilization above 85 to 90 percent of a circuit's capacity degrades everything on it regardless of what is normal for the season. Voice-grade lines are fixed truths anchored in standards and practice; our VoIP monitoring guide covers the specific latency, jitter, and loss ceilings. License and contract boundaries are fixed truths. Two practices keep static thresholds useful. Alert on sustained breach rather than instantaneous spikes, because momentary peaks are what networks do, and a duration condition (five minutes above threshold, for instance) filters the noise without hiding the events. And set thresholds per object rather than globally, because a 10 Gbps core link and a 50 Mbps branch circuit share no meaningful fixed values except percentage of themselves. ## Baseline deviation: Where it belongs Everything without a fixed truth needs a learned one. Host outbound volume, DNS proportions, conversation fan-out, per-site application mix: normal differs per object and per hour, and the meaningful signal is deviation from that object's own history. Baseline-driven alerting catches the security patterns (beaconing, lateral movement, exfiltration) and the operational anomalies (the backup that started running at noon) that no static number can express. Baselines need feeding and maintenance: Two to four weeks of history before trusting them, separate populations in separate groups (servers, workstations, scanners, backup infrastructure), and scheduled review as the network changes. A baseline built across mixed populations averages them into a normal that describes none of them. ## Severity tiers that map to response Severity means nothing unless each tier maps to a distinct response. A three-tier scheme covers most environments: 1. **Critical, pages a human now:** Conditions causing active user impact or active security risk. Saturated primary links, exfiltration-pattern matches, voice thresholds breached on production paths. 2. **Warning, lands in the queue:** Conditions trending toward impact. Utilization climbing into the 70s on a growth curve, baseline deviations without a security shape, single-threshold voice degradation on secondary paths. 3. **Informational, appears in review:** Conditions worth a look during working hours. First-seen destinations without corroborating signals, threshold touches that self-resolved. The discipline is holding that line. An alert that wakes someone at 3am must come with a 3am action attached, and everything else waits for morning because you decided it should. ## The alert fatigue math Fatigue follows arithmetic. A rule that fires ten times a day at 95 percent false positive rate trains its audience in under a week, and the training persists after the rule is fixed. Track two numbers per rule: fire rate and action rate (how often the alert led to any action). Rules with high fire rates and near-zero action rates are fatigue generators, and the fix is tuning them, demoting them a tier, or deleting them. Deleting a useless alert is an improvement, and teams that cannot delete alerts accumulate them until the pager becomes background noise. ## The review loop Alerting decays without maintenance, because networks change and thresholds do not change themselves. A monthly review covering four questions keeps the system honest: 1. Which rules fired most, and what fraction led to action? 2. Which incidents arrived without an alert, and what rule would have caught them? 3. Which baselines cover populations that changed (new sites, migrated services, retired hosts)? 4. Which thresholds were set during an incident and never revisited? An hour a month on these four questions is the cheapest reliability investment most teams can make. ## Alerting with ManageEngine NetFlow Analyzer ManageEngine NetFlow Analyzer implements both threshold models on the same flow data: static thresholds on utilization, volume, and voice metrics, and behavior-based alerting against learned baselines. ### Feature highlights: - **Per-object thresholds:** Utilization and volume conditions per interface, group, and site. - **Sustained-breach conditions:** Duration qualifiers that filter momentary spikes. - **Severity and routing:** Tiers mapped to notification channels, email, chat, and ticketing integrations. - **Alarm context:** Each alert links to the conversations behind it, which shortens the action decision. ## FAQs on traffic alerting ### What utilization threshold should I set on a link? Alert on sustained utilization in the 85 to 90 percent range for most links, with a duration condition of several minutes to filter momentary peaks. Set warning-tier alerts lower, around 70 to 75 percent, to catch growth trends before they become saturation events. ### How long before baselines are trustworthy? Two to four weeks of history in most environments, with separate baselines per population. Trusting a baseline earlier produces false confidence, and mixing populations produces a normal that describes none of them.