Hardcraft

Back to services

Monitoring and alerting

Knowing something broke before your customer calls — and only being woken when it actually matters.

Most monitoring fails in one of two ways: it measures nothing of value, or it cries wolf so often that nobody looks any more. Both are as bad as having none.

I set up what actually says something about your service:

  • Availability and response time of the application, not just ping
  • Disk, memory, CPU and I/O with thresholds that match your load
  • Certificates, domains and expiring keys
  • Backup results — a failed backup is an incident
  • Alerting through the channel you actually watch

Usually built on Prometheus with Grafana, or Zabbix where that fits better with what is already in place.

Is your infrastructure running the way it should?

Half an hour on the phone costs nothing and usually surfaces a few concrete improvements already.

Get in touch