Running services, network usage, memory usage, bandwidth, disk I/O, successful logins, whether the thing is even alive, etc…
So far my only method has been “hope and pray”.
Gatus. Works and easy incremental set up.
Submitted 22 hours ago by user224@lemmy.sdf.org to selfhosted@lemmy.world
Running services, network usage, memory usage, bandwidth, disk I/O, successful logins, whether the thing is even alive, etc…
So far my only method has been “hope and pray”.
Gatus. Works and easy incremental set up.
Use service, wait for 404 or data loss…
You guys are monitoring your servers?
“can’t access this thing, server must be down”
I’ll fix it when I get home
I plan to perhaps, just maybe, I am not sure, to try self-hosting an email server. But I’ll need a good uptime.
Although I am really scared about security.
I self-host my own email, have done so for years. I’ve not gotten pwned. Use an SSH key and a firewall, only expose the ports that should be public, have reasonable access controls, etc, just practise general cybersecurity—it applies to servers too.
I know everyone complains about hosting email. But I’ve not run into major problems. My emails get through spam filters and I have good uptime. It’s also a fun learning experience.
Been running one for many years, mostly issue-free. Had to request whitelists way in the beginning, but things turned out ok.
Check out this super great guide. workaround.org/ispmail-trixie
Well, it’s a good learning experience. But…… don’t. Good luck getting around the major email providers’ smart screen shit. It’s probably not gonna happen. I tried a few years ago, and even microslop couldn’t figure out why their shit screen crap wouldn’t allow my emails. But it’s worth doing it in the short term to learn how to do it.
Beszel & Pangolin.
I notice something stopped working, or someone in the household notices.
Then I go fix it.
And the best time to notice that the backups weren’t working is when you need access to the backup. 😅
Backups are pretty much the only thing I get notifications for. Otherwise, one app or another on my phone will complain if it loses connection to my server.
I have a robust monitoring system for my Jellyfin server, been running for the last few years. Checks in periodically, at least once per 24 hours and notifies me if it’s down. Doesn’t use any electricity, but does consume a good amount of Cheerios and mac & cheese.
Lies! It uses electricity!
Indirectly. But as another commenter pointed out, it balances out with the tax rebates.
Icinga and Phare Uptime
Ultime Kuma and Beszel
That’s the neat part, I don’t.
Lights are on, server is on.
I recently found out that this is not always true.
I never said anything about it responding.
Kernel panic?
My kernel is chill.
Monit, simple to setup, gets the job done, runs on my router too.
I have this on my NAS, works well
Whining from end users.
It’s running the private DNS for my phone. If it goes down, I realise quite quickly.
I go the easy route with munin and wazuh. It comes with a ton of stuff preconfigured
I pay a group of schoolchildren to refresh a group of browser tabs and yell if anything isn’t responding.
There are plenty of previous threads with recent answers if you search.
Well I used to use VGA for my monitors, now its mostly displayport.
Seriously though, prom+Loki+alloy+a few other things to Grafana, with alerting in grafana.
I am using prometheus-nodeexporter which is scraped by Prometheus and then visualized in Grafana
Uptime Kuma! For HDDs I have them report their “ping” as percentage full
Can you teach me this wizardry?
Sure! In Uptime Kuma, add a monitor with type Push. It gives you a URL with a unique token, and the monitor goes into “down” state if nothing calls that URL within the heartbeat interval. Then have a cron (or a systemd timer or whatever) run this on the machine with the drive every 5 or 10 minutes:
use=$(df --output=pcent /mnt/yourdrive | tail -1 | tr -dc ‘0-9’) curl -fsS -G “http://your-kuma:3001/api/push/TOKEN --data-urlencode “status=up” --data-urlencode “msg=${use}% used” --data-urlencode “ping=${use}”
Ping is of course meant to be response time, but Kuma will graph any number you toss in, so you get a “filled” chart for the drive. I also send status=down once usage passes a threshold. And if the box itself should fall over for some reason, the missing check-ins trigger the alert too.
(I typed this quick on my phone, so please double-check and probably don’t blindly copy-paste)
Grafana dashboard
image You can find more details and the source code here: https://erasmus.works
Ooooo, that looks nice.
Acronyms, initialisms, abbreviations, contractions, and other phrases which expand to something larger, that I’ve seen in this thread:
| Fewer Letters | More Letters |
|---|---|
| AP | WiFi Access Point |
| DNS | Domain Name Service/System |
| LXC | Linux Containers |
[Thread #122 for this comm, first seen 9th Oct 2026, 13:00] [FAQ] [Full list] [Contact] [Source code]
Uptime Kuma for outages, Prometheus for metrics, rsyslog to Loki backend for logs. Grafana for ingesting Prometheus and Loki data.
Cockpit. Total gamechanger for me. I don’t have to ssh to my boxes individually to see what’s going on. Instead it’s a web UI AND you can connect basically every server you want to one instance.
github.com/cockpit-project/cockpit
That said. For individual servers?
htop/nmon (nmon is FANTASTIC)
ps -ef | grep <whatever you’re looking for>
dmseg
tail -f /var/log (or whatever log file you care about)
netstat
those are all that come to mind right away.
It’s amazing the risk profile of that.
How do you configure adding multiple servers to the same instance?
So I actually just went to make a guide on this, demonstrating on my ubuntu 26.04 lab vps, and it appears the connect option is not there.
However on my 24.04 lab vps, it’s there. Same with my local 22.04 systems (yes I know I have to upgrade lmao)
I went to cockpit docks and sure enough, the feature is deprecated :( Dang.
docs.cockpit-project.org/…/feature-machines.html
It was really cool though while it worked.
Lol I go to all of my server’s specific cockpit addresses, sounds like I’m doing it wrong, I’ll need to look at connecting them
So it turns out… you can’t connect them anymore… The feature is deprecated
docs.cockpit-project.org/…/feature-machines.html
(also, I wanted to just link my response from another user who had asimilar question but for whateVER reasyon, Alx won’t let me??)
My 3-node Proxmox cluster has a dashboard, my NAS has a dashboard, and Home Assistant has a dashboard. Hell, my gaming PC has a dashboard, too (Webmin and Cockpit).
I use Alloy to collect Metrics of the host (Disk usage, CPU, Ram, etc.) and different logs. With Grafana the data is then displayed as a daschboard. If everything goes south, Alertmanager sends Mails to me
Uptime Kuma for monitoring web service status (just like if the jellyfin, etc. website is up), and Zabbix for monitoring system metrics like CPU/Memory/Disk I/O, etc. Zabbix can do a whole lot, but can also get a little complicated. There are plenty of guides out there for configuring it though.
I just look at the console.
I get a text pretty quickly.
I usually just walk down the hall and move the mouse. I don’t really need remote monitoring. 😅
Netdata, looking at dashboards when something breaks :)
timochka@lemmy.zip 1 hour ago
I use Signoz and OTel collectors to bring in everything from Kubernetes (and also SNMP-to-OTel for the network and NAS stats.)
That covers dashboards and basic alerting. Then on top of that I have an n8n workflow that runs every morning - it pulls all the main stats out of Signoz, as well as issues and their status from my infrastructure repo in Gitlab, and feeds it all into a local LLM (Qwen-3.8-Flash-Next at the moment), which emails me a report on overall cluster health. It also automatically opens a ticket in Gitlab for anything that might need fixing.
It works really nicely, gotta say.