Using Statistics and Sensors to Ensure Network Availability
Using Statistics and Sensors to Ensure Network Availability
Turn device health, traffic statistics, logs, and alerts into early warning signals. Learn to establish normal performance, recognize anomalies, and select the right monitoring evidence before users experience an outage.
Availability Begins Before the Failure
Network availability is the ability of systems and services to remain reachable and usable when required. Downtime interrupts operations, weakens user trust, and can damage revenue and reputation. Effective operations therefore emphasize proactive monitoring: observing health and performance continuously so warning signs are discovered before they become outages.
Observe
Collect metrics, events, logs, and traffic summaries from devices and services.
Compare
Compare current behavior with a baseline, threshold, SLA, or expected configuration.
Act
Alert the right team with enough context to investigate and correct the condition.
Monitoring is essentially listening to what the network is reporting. A useful system helps administrators detect early failure indicators, recognize degradation, establish normal behavior, and respond using actionable evidence.
Device and Chassis Health
Hardware and operating systems expose counters much like a vehicle dashboard. A single high reading may be temporary; a sustained or correlated pattern is more significant.
Temperature
Heat rises with blocked airflow, dust, fan failure, or high load. Excessive temperature can cause throttling, shutdowns, reboot loops, and permanent damage.
CPU utilization
Sustained processor use above roughly 85% suggests overload. High user time points to application work; high interrupt time may indicate hardware or driver trouble.
Memory utilization
Committed memory above roughly 80%, available memory below 5%, or heavy paging indicates pressure. Expanding memory pools can reveal leaks.
| Counter or sensor | What it reveals | Warning clue from source |
|---|---|---|
| % Processor Time | Time executing non-idle work | Sustained above 85% |
| % Interrupt Time | Time servicing hardware interrupts | Unexpectedly high values |
| Processor Queue Length | Threads waiting for CPU | More than twice the CPU count |
| Committed Bytes in Use | Committed virtual memory pressure | Above 80% |
| Available Memory | Unused physical RAM | Below 5% is critical |
| Paging activity | Disk use caused by insufficient RAM | Persistent or excessive activity |
How Well Data Moves
Bandwidth vs throughput
Bandwidth is a link's theoretical capacity in bits per second. Throughput is the actual rate successfully transferred. Sustained use near 70% may suggest saturation depending on the technology and traffic pattern.
Output queue
An output queue forms when frames wait to leave an interface. A persistent queue length above 2 in the source's example indicates backlog and possible congestion.
Latency
The end-to-end delay, often measured as round-trip time with ping. Long routes, overloaded links, inspection devices, and processing delays can increase it.
Jitter
Variation in latency from packet to packet. Voice, video, and streaming need predictable arrival times, so jitter may hurt quality even when average latency seems acceptable.
Packet loss should be near zero on a healthy internal network; real-time traffic may not retransmit what disappears.
Normal First, Anomaly Second
A baseline records normal performance under representative conditions. Without one, a value such as 65% link utilization lacks context: it may be routine for that link or a serious deviation.
Measure during normal operation, including relevant busy and quiet periods.
Collect over enough time to reveal daily, weekly, seasonal, and workload patterns.
Update the baseline after upgrades, migrations, topology changes, or major workload shifts.
Alert when current behavior deviates materially from the expected range.
An anomaly is behavior that differs from what is standard, normal, or expected. Automated anomaly alerts reduce manual inspection and shorten response time, but poorly tuned thresholds create alert fatigue.
Discovery, Performance, Availability, and Configuration
| Monitoring function | Primary question | Typical outcome |
|---|---|---|
| Network discovery | What devices exist and how are they related? | Inventory, topology, rogue-device detection |
| Performance | How are resources and traffic behaving? | Trends, bottlenecks, capacity planning |
| Availability | Is the device or service responding? | Uptime, downtime, SLA evidence, outage alert |
| Configuration | Does the device match approved settings? | Drift detection, compliance, change evidence |
Ad hoc discovery is manual and on demand. Scheduled discovery runs repeatedly and is better suited to finding new, missing, or unauthorized devices. Performance platforms store historical counters for forecasting, while availability checks measure service responsiveness rather than simply whether a device has power.
Polling Devices and Receiving Traps
Simple Network Management Protocol (SNMP) allows a central platform to retrieve and organize management data from network devices.
Agent and NMS
The agent runs on the managed device. The Network Management Station polls agents, stores results, displays status, and receives alerts.
OID and MIB
An Object Identifier uniquely addresses a managed object or counter. The Management Information Base defines and organizes those OIDs hierarchically.
Get, Set, and Trap
Get retrieves data; Set changes a parameter and is used cautiously; a Trap is an unsolicited agent alert.
Ports
UDP 161 carries Get/Set exchanges. UDP 162 carries traps and informs toward the NMS.
| Version | Security and capability | Use |
|---|---|---|
| SNMPv1 | Legacy and insecure | Avoid where possible |
| SNMPv2c | Improved counters and GetBulk; community strings remain plaintext | Common legacy monitoring |
| SNMPv3 | User-based authentication, integrity, encryption, and access control | Preferred secure version |
Packets, Mirrors, and Flows
Packet capture
Wireshark and similar analyzers capture raw frames and decode protocols across layers. Packet detail is excellent for deep troubleshooting but creates large datasets.
Port mirroring
A switch copies selected traffic to a monitoring port. SPAN is local, RSPAN extends mirroring remotely at Layer 2, and ERSPAN transports mirrored traffic across a routed network.
Flow data
NetFlow summarizes conversations instead of storing every packet. It identifies top talkers, common destinations, protocols, usage patterns, and anomalies.
| Evidence | Detail level | Typical fields or requirement | Best suited to |
|---|---|---|---|
| Packet capture | Full frame/packet detail | Promiscuous NIC and visibility through TAP or mirror | Protocol errors and exact payload/sequence behavior |
| NetFlow | Conversation summary | Source/destination IP, ports, protocol, interface, ToS | Top talkers, trends, capacity, anomaly hunting |
| SNMP counters | Device/interface statistics | OIDs from agent | Health, utilization, errors, status |
Centralizing Events with Syslog
Application, system, security, traffic, and audit logs record what systems observed or did. A central collector preserves evidence even if a source device fails or an attacker alters local history, and it allows events from many systems to be correlated.
Syslog commonly sends messages using UDP 514. Traditional UDP delivery is fire-and-forget, so delivery is not guaranteed. Implementations may support more reliable or protected transports, but remember UDP 514 for the exam.
| Level | Name | Meaning |
|---|---|---|
| 0 | Emergency | System unusable |
| 1 | Alert | Immediate action required |
| 2 | Critical | Critical condition |
| 3 | Error | Component or operation failure |
| 4 | Warning | Warning condition |
| 5 | Notice | Normal but significant event |
| 6 | Informational | Informational message |
| 7 | Debug | Detailed troubleshooting output |
From Collected Logs to Security Intelligence
A Security Information and Event Management (SIEM) platform extends log aggregation with parsing, normalization, correlation, dashboards, detection rules, alerts, investigation, retention, and compliance reporting.
SEM
Security Event Management emphasizes real-time monitoring, correlation, and response to current events.
SIM
Security Information Management emphasizes storage, search, reporting, and historical analysis.
A SIEM can connect an authentication failure on one server, firewall traffic from the same address, and an endpoint alert into a single incident. It does not guarantee correct conclusions: useful detections still require reliable time synchronization, normalized data, tuned rules, context, and human investigation.
Choose the Best Monitoring Evidence
Select a scenario to reveal the most useful starting metric or monitoring source.
The recommended evidence will appear here.