Skip to content

Who this is for

For companies where downtime costs more than maintenance, and the in-house IT team has more urgent work than sitting on server duty.

What you get

  • Monitoring in Zabbix with Grafana dashboards and thresholds tuned to real traffic.
  • Incident response within an agreed time, with communication while the incident is ongoing.
  • High-availability clusters: Percona XtraDB Cluster, HAProxy, replication and failover.
  • Performance tuning: SQL queries, caching and application server configuration.
  • Monthly security updates and configuration reviews.
  • Backup verification through test restores, not through a job status page.
  • A report on work done and environment health that makes sense outside the IT department.

Problems I typically solve

  • You learn about outages from customers, not from monitoring.
  • Alerts fire so often that nobody reads them any more.
  • Performance degrades month over month and nobody knows why.
  • One person knows the configuration and is about to go on holiday.
  • Every update is a risk, so systems have not been patched in a year.

Stack involved

Monitoring & logging

  • Zabbix
  • Grafana
  • New Relic
  • Percona Monitoring & Management
  • Rsyslog
  • Graylog

Databases & queues

  • Percona XtraDB Cluster
  • MySQL
  • RabbitMQ

Systems & services

  • Ubuntu
  • Debian
  • Postfix
  • Samba
  • NFS
  • CUPS
  • Rocket.Chat
  • PHP / Magento

How I work

Four steps, always in this order

No audit means no design, no design means no rollout. This order saves money by the third step.

  1. Step 01

    Audit and analysis

    An inventory of what is actually running and a risk list ordered by business impact. Without this step everything that follows is guesswork.

  2. Step 02

    Architecture design

    The target shape of the environment, with cost, schedule and a rollback plan. Decisions are made before the rollout, not during it.

  3. Step 03

    Implementation

    Staged delivery, with configuration described in Ansible and Terraform. Every step is repeatable and reversible.

  4. Step 04

    Operations and monitoring

    Monitoring with thresholds tuned to real traffic, updates, restore testing and a report on the state of the environment.

Frequently asked questions

Contact

Describe the problem and get a specific answer

The fastest way to get to the point is to outline your current environment and what needs to change in your first message.

················