Jump to content

Recommended Posts

Posted

Good morning,

 

Do any of you have software setup that alerts you if the internet is down, your server reboots, or other network health info?

 

I need something to alert me when I’m not there. 
 

Thankyou. 

Posted

PRTG (up to 100 sensors are free) and The Dude (you need the old 4.0 beta version of this, otherwise it requires Mikrotik hardware). The former is more for the away-from-site monitoring, and the latter we have up on a screen in the office as we also use it to keep an eye on LAN and internet load.

  • Like 1
Posted

We have NetXMS, which is free. It can send email/telegram/teams message for alerts.

 

Also have all the "Entra ADDS Health connect" stuff installed & Entra will automatically send alerts if it detects a problem with your AD, including internet outages (Sync failures)

  • Like 1
Posted

CheckMK is pretty good. It is highly customisable. The major downside is the lack of mobile interface.

  • Like 1
Posted (edited)

+1 for Checkmk. The Community edition is the free self-hosted option. Running the Windows agent on your servers opens up a whole world of usefulness.

Edited by jthompson
  • Like 1
Posted

We use PRTG free, to keep an eye on switches, internet, internet speeds, copiers, and main servers are up.  Hoping to have some time (ha ha ha) to move it all to something else, so we can monitor more, as the 100 free devices is restricting us now.

  • Like 1
Posted (edited)

I have used PRTG in the past but the only thing you will find is that you may quickly use up 100 sensors once you see what you can do with it. But prob one of the easy monitoring tools to setup but I think it only runs on Windows??

 

Zabbix can run as a docker so pretty easy to get up and running, I'd always setup a linux VM specifically for the docker though and not have it running direct on hardware. 

 

At home I use a combination of Zabbix and Uptime Kuma with notifications via home assistant and Discord. Uptime Kuma can do DNS lookups, https cert checks, website checks, SMB etc. and alert if any fail.

 

 

Edited by Davit2005
  • Like 1
Posted
16 minutes ago, Davit2005 said:

I have used PRTG in the past but the only thing you will find is that you may quickly use up 100 sensors once you see what you can do with it.

 

 

100 ? We are running 788,000 metrics across the entire infrastructure!

  • Like 1
Posted
15 hours ago, mrbios said:

Zabbix, screenshot of my dashboard attached as an example of what I'm monitoring. Screenshot taken out of hours so all the figures are low, surprised there's still 80 live users on sophos at nearly 6pm though lol

 

The honeycomb highlights red if uptime goes over 1 month (suggesting windows update automation hasn't gone ahead or delayed). Orange for uptime under 12 hours so that anything that's rebooted recently is highlighted. Zabbix agent is installed on all VMs and hosts, flags up any services that are set to Automatic startup if they're stopped etc.

zabbixdash.png

 

I stick to the LTS version, 8.0 should be out soon hopefully, looking forward to new features.

 

Sorry I am shamelessly stealing some of these widgets :) we already run zabbix in house for a while now so I have the data already. Plus we also have Sophos so the only words I can say are ZOINK and Thanks for sharing ! 👍

  • Like 1
Posted

We use Mutiny from mutiny.com it's simple and easy to use/configure and has saved our bacon so many times as we've either been able to resolve issues before they become one (like dangerously low disk space on the C :  drive) or spotting issues as they happen before our Academies got around to reporting it to us like we can see if an Academy has likely had a power cut or their internet connection has gone down.


When people say you need to "be more proactive" Mutiny is the only system I've ever used that has made us be more proactive as you can see when a server/device is starting to get low on disk space or when it's using up all its allocated CPU cores or RAM and you can do something about it.

  • Like 1
Posted

i use librenms purely for automated alerts, I dont use the dashboard unless im chasing alerts.  The alerts are a bit clunky to setup to begin with, but work well enough and have a decent logic path.  I have it set up for disk space, network saturation on the switches, server status (memory pressure, cpu levels sustained over x time etc), port flaps, trapping RSTP events, the usual stuff.  It is SNMP so fairly basic if you want detail and data as you are limited to what the devices or agents expose.  Still, you can still get a lot of SNMP data from your monitored devices and as long as you are happy to add various logic strings to your alert monitors then it works well.  The cost is your time.  

 

Any type of monitoring can get quite interesting, especially when you can drill down to your switch port throughputs and see what culprits use the most bandwidth.  I was quite surprised by how chatty some clients are (naturally, without being misconfigured) and how efficient others can be (I was surprised how efficient hikvision CCTV cameras are compared to verkada).

  • Like 1
Posted
On 02/06/2026 at 08:27, dmj said:

100 ? We are running 788,000 metrics across the entire infrastructure!

Soz was referring to the free version, lol.

Posted

100 sensors is not a lot.  That wont cover more than a couple of VMs.  I havent counted the number of sensors I use, one switch stack would be about 2000 I guess, multiply that out for a few switch stacks, then buildings, then sites.  It adds up really quickly.  I think our UPS eats about 40 alone with each battery bank having temperature, voltage, fans, load, runtime. 

 

Sure, if you wanted a basic "Im happy that the UPS is working" then you could only use a single sensor for "temperature" I guess, but I would probably want one for each battery bank so thats 6.  Then for each VM maybe monitor disk space alone on one drive and have fancy logic for up/down.  For switches it would be pointless really, they expose a lot of data.  Printers are odd, the toshibas we use need a bit of logic across a few sensors as I only care about certain bits that papercut doesnt tell me (staples, waste toner and jams) but need a few sensors each.

Posted

You have to be selective about what you monitor, but it works for us.

 

The reverse question is how often are you being alerted or looking at network speeds/traffic on every single port of your switches?  We have ping to each switch, to know they are up, and the external traffic port for the 4 sites to keep an eye on anomalies. We aren't monitoring every disk drive, but we have very few now, and even less that fluctuate wildly.

 

Copiers we have set to let us know when toner reaches 0 to prompt a tech to go replace, and other alerts we pass straight to the copier team to service next time they are out, but that is all counted as one sensor, per device.

 

Other sensors I have monitor number of admins in the server and local pc groups, as a canary in case of an attack, M365 services, alarm panel responsiveness (as we had a spate of them giving up communicating), certificate expiry, Backup status and the last of the CCTV DVRs. So pretty well covered, without being overwhelmed or fretting about every little blip.  I do wish I could monitor network inter-links, to be able to pinpoint more easily where traffic spikes are coming from, but that is why I am looking at changing.  PRTG has been fantastic to get something in, easily, and introduce monitoring to the team.

 

Part of me would like to monitor all the things but I also know that the vast majority of the data would never be used or looked at, so why do it?

 

It is about being proactive, we normally inform the Estates team that the electricians are running wild or that the power has gone out in areas before they know.

  • Like 1
Posted

This might be a laugh, but we just rely on an email from Microsoft called "Password Hash Synchronization heartbeat was skipped in last 120 minutes"  which triggers us to check our servers instantly. A very lazy solution I know and do need to look at something more robust!

  • Like 2
Posted

As often as a person loops back a cable.  Nice to know it has been done as this doesn't affect the day to day running (as RTSP sorts it out) but you dont necessarily know it has been done.  If students are messing about then other stuff has most likely happened in that class, keys being switched on keyboards etc.  So I would rather get an alert for that, get it passed over to the behavioural team so that they can sort it out in a class.  Ive also had a port flapping issue on a base unit picked up a few times, these stone AIO had NIC and WIFI, the NIC was misbehaving and flapping with the WIFI.  Ive had a misconfigured AIO that was flapping and had internet sharing enabled, this naturally looped back with RTSP doing its nut, the AIO worked just fine but was a bit slow.  Misconfigured software filling up logs on C:, UPS with a faulty fan.  idrac talks nicely, ive had a faulty drive pick up via idrac forgetting that I also had a machine agent that did the same.   I do have logic for our >10Gb links operating at 75% over a period of greater than 1 minute, if im hitting that level on my interconnects then there is an issue for us.  Getting historical data on hypervisor and backupserver nic loads pinpointed a misconfiguration with SET (teaming) when everything appeared to be fine.

 

Ironically enough, I dont bother with toner as the papercut alert worked so I didnt bother setting a sensor up for that. 

Posted

Graphing can sometimes illustrate a situation really well, where you'd otherwise you'd be needing collate bits of information from here and there. It can also help you get a sense of how something is trending over time.

Printer toner level graphing can give you a really quick sense of whether a printer is being overused or underused at different times of year, how it's usage compares to others in the fleet, etc. That information may be available in other forms, but decent graphing can surface it really easily and quickly. Similarly for graphing total printed pages for a printer, disk/drive usage, device uptime, WAN traffic: interactive graphing can visualise a situation for you really easily, in a way that threshold alerts or alarms alone perhaps can't.

 

Also, everyone knows that a dashboard with graphs in can be dead pretty.

  • Like 2
Posted
On 03/06/2026 at 11:56, jthompson said:

Also, everyone knows that a dashboard with graphs in can be dead pretty.

We used to have loads across our office wall. It's all only online now I mostly work from home :(

Posted
On 03/06/2026 at 11:56, jthompson said:

Graphing can sometimes illustrate a situation really well, where you'd otherwise you'd be needing collate bits of information from here and there. It can also help you get a sense of how something is trending over time.

Printer toner level graphing can give you a really quick sense of whether a printer is being overused or underused at different times of year, how it's usage compares to others in the fleet, etc. That information may be available in other forms, but decent graphing can surface it really easily and quickly. Similarly for graphing total printed pages for a printer, disk/drive usage, device uptime, WAN traffic: interactive graphing can visualise a situation for you really easily, in a way that threshold alerts or alarms alone perhaps can't.

 

Also, everyone knows that a dashboard with graphs in can be dead pretty.

 

Graphing is also very C-level friendly when negotiating project budgets.

Create an account or sign in to comment

You need to be a member in order to leave a comment

Create an account

Sign up for a new account in our community. It's easy!

Register a new account

Sign in

Already have an account? Sign in here.

Sign In Now



×
×
  • Create New...