Small business network monitoring comes down to watching five things: whether each circuit is up, how good the connection is (latency, jitter and packet loss), how full it is, whether your firewall, switches and Wi-Fi are healthy, and what the security logs say. Each needs a sensible alert that reaches a named person, and at least one monitor has to sit outside your office, because a monitor inside a dead office cannot tell anyone.
Monitoring sounds like an enterprise discipline with wall-sized dashboards. For a small business it can be much simpler, and the payoff is concrete: you hear about problems before customers do, you can prove to a carrier when and how a circuit failed, and you can see when an upgrade is due rather than guessing.
Why it is worth the trouble
- You find out first. A failed backup circuit or a dying switch is invisible until the day you need it. Monitoring finds silent failures.
- Carrier tickets go faster. "The internet is slow" gets a script. "Packet loss to your gateway rose to 4% at 10:12 and is still there" gets an engineer.
- SLA claims need timestamps. Service credits usually depend on when an outage started and ended, and on you reporting it promptly. Our guide to what a 99.99% SLA buys explains why the clock matters.
- Upgrades get planned. Utilisation trends tell you when a circuit is filling up months before staff start complaining.
The five things worth watching
| What | Why it matters | How it is usually measured | A reasonable starting alert |
|---|---|---|---|
| Circuit up or down | The basic question, for primary and backup | Regular test traffic to the circuit's gateway and to outside targets | Several missed checks in a row, not one |
| Connection quality | Calls and video fail on loss and jitter long before a circuit is "down" | Continuous tests measuring latency, jitter and loss | Loss or jitter above your normal baseline for several minutes |
| Utilisation | A full circuit feels like a broken one | Interface counters polled from the firewall or router | Sustained high use at peak, reviewed weekly rather than paged |
| Device health | Firewalls, switches and access points fail too | Polling of CPU, memory, temperature, power and fan status | Device unreachable, or a fault such as a failed power supply |
| Security and expiries | Attacks, blocked threats, lapsed licences and certificates | Logs sent from the firewall; date tracking for renewals | Repeated failed logins, or a renewal date 30 days out |
Start with those thresholds and tune them. The goal is alerts that mean something, not a flood that everyone learns to ignore.
How the tools collect the data
Simple reachability tests
The oldest method is still useful: send a small test packet to a device or address at regular intervals and record whether it answers and how long it took. It shows up, down and latency, and with enough samples, packet loss and jitter.
SNMP polling
Most business network equipment supports SNMP, a standard way for a monitoring system to ask a device for its counters: traffic in and out of each port, CPU and memory, errors, temperature. Polling every few minutes builds the utilisation graphs. Use the newer, authenticated version of the protocol where your equipment supports it, and never leave default community strings in place.
Flow data
Firewalls and routers can export flow records (NetFlow, IPFIX or sFlow, depending on the vendor) that describe which devices talked to which destinations and how much data moved. This is how you answer "what is using all the bandwidth?" without guessing.
Logs
Devices can send event logs to a central collector: logins, configuration changes, blocked connections, interface flaps. Logs are where security problems and the cause of an outage usually show up.
Synthetic tests
Some tools simulate what users do: load a web page, resolve a domain name, or measure the quality a voice call would have. These catch problems that device counters miss, such as a DNS failure that leaves the circuit up but the internet unusable.
Cloud-managed dashboards
Many current firewalls, switches and access points report to the vendor's cloud dashboard automatically. For a small office this may cover most of the list above with very little setup. Check what it alerts on by default, and who it alerts.
Watch from outside, too
A monitoring system inside your office shares the fate of your office. If the only circuit fails, the monitor cannot send its alert, and you find out when someone walks in. Two fixes are simple:
- An external check. A monitoring service on the internet tests your office's public IP address or a device behind it. When it stops answering, the alert comes from outside.
- An out-of-band path. A small cellular modem or the backup circuit carries alerts and remote access when the primary is down.
External tests also give you an independent record of outages, which is exactly what you want in front of you when a carrier disputes the timeline.
Alerts without the noise
- Require persistence. Alert after several consecutive failures or a few minutes of degradation, not one lost packet.
- Separate urgent from informational. A circuit down during opening hours is a phone call. A circuit at high utilisation is a weekly report.
- Name the person. Every alert goes to someone specific, with a backup person when they are away. "The IT inbox" is not a person.
- Write the first step. A short note with each alert type: who to call, the circuit ID, the carrier's support number.
- Mute planned work. Set maintenance windows so a scheduled firmware update does not wake anyone.
Reading utilisation correctly
Utilisation graphs are averages, and averages hide peaks. A circuit that averages 30% over a five-minute sample might have spent part of those minutes completely full. Look at the busiest hours of the busiest days, not the daily average. A common rule of thumb is to start planning an upgrade when a circuit runs at a high share of its capacity, often quoted at around 70 to 80%, for sustained periods at peak. Some carriers bill burstable services on the 95th percentile, which ignores the busiest 5% of samples; if yours does, your graphs should use the same method. Our guide to how much bandwidth you need covers the sizing side.
A worked example
A hypothetical illustration. The office and figures are invented.
A 25-person office has a 300 Mbps fiber primary and a cable backup. The owner sets up a cloud dashboard for the firewall and access points, an external check on the office's public IP every minute, and continuous quality tests to two outside targets.
Three things happen in the first quarter. In week two, the dashboard shows the backup circuit has been down for eleven days; nobody noticed because the primary was fine. It is fixed before it is needed. In month two, quality tests show packet loss climbing to around 2% most afternoons while utilisation stays modest, so the cause is not a full circuit. The carrier gets a ticket with times and graphs, finds a faulty port, and replaces it. In month three, utilisation graphs show the circuit near its ceiling every day from 9 to 10am when overnight backups overrun into the morning. Moving the backup schedule fixes it without an upgrade. None of this needed a network engineer on staff, but it did need someone to look.
Do it yourself or have it managed
A single office with cloud-managed equipment and an external uptime check can often be monitored in-house, provided someone owns the alerts. Once you have several sites, more than one circuit per site, or voice systems that need quality monitoring, a managed service is often cheaper than the staff time. Our managed network service covers monitoring, alerting and the carrier escalation that follows. Security logging is its own discipline; see our cybersecurity page.
A monitoring checklist
- Primary and backup circuits each tested independently.
- At least one external check on the office.
- Quality tests for latency, jitter and loss, if you run voice or video.
- Utilisation graphs kept for at least a year, so you can see trends.
- Firewall, switches, access points and battery backups reporting health.
- Firewall logs sent somewhere other than the firewall.
- Licence, certificate and contract renewal dates tracked.
- Each alert routed to a named person with a written first step.
When calls are the problem, our guide to VoIP call quality troubleshooting shows how monitoring data narrows the cause, and latency, jitter and packet loss explained covers what the numbers mean.
Questions to ask your provider
- Do you monitor my circuit proactively, and will you contact me when it fails, or do I have to open a ticket?
- Can I see utilisation and performance data for my circuit?
- How do I report an outage so the SLA clock starts, and what information do you need?
- Does my firewall or router support SNMP and flow export, and are they configured?
- If you offer managed monitoring, what is watched, who is alerted, and at what hours?
- Is the backup circuit monitored separately from the primary?
Getting circuits worth monitoring
Monitoring tells you when a circuit is the problem. Fixing that sometimes means a better circuit or a better backup. We qualify your address across several Tier 1 carriers and bring back the best offer, with no markup because the carrier pays us, and one agent, Phil Morales, who handles the ticket when monitoring flags something. The quote is free with no obligation, and we usually respond the same day. Send us your address, call 478-758-8091 or text (347) 870-0965.