Managed operations (SLA)
I take on day-to-day responsibility for keeping the environment running: monitoring with meaningful alert thresholds, incident response and steady removal of root causes.
Request a quoteWho this is for
For companies where downtime costs more than maintenance, and the in-house IT team has more urgent work than sitting on server duty.
What you get
- Monitoring in Zabbix with Grafana dashboards and thresholds tuned to real traffic.
- Incident response within an agreed time, with communication while the incident is ongoing.
- High-availability clusters: Percona XtraDB Cluster, HAProxy, replication and failover.
- Performance tuning: SQL queries, caching and application server configuration.
- Monthly security updates and configuration reviews.
- Backup verification through test restores, not through a job status page.
- A report on work done and environment health that makes sense outside the IT department.
Problems I typically solve
- You learn about outages from customers, not from monitoring.
- Alerts fire so often that nobody reads them any more.
- Performance degrades month over month and nobody knows why.
- One person knows the configuration and is about to go on holiday.
- Every update is a risk, so systems have not been patched in a year.
Stack involved
Monitoring & logging
- Zabbix
- Grafana
- New Relic
- Percona Monitoring & Management
- Rsyslog
- Graylog
Databases & queues
- Percona XtraDB Cluster
- MySQL
- RabbitMQ
Systems & services
- Ubuntu
- Debian
- Postfix
- Samba
- NFS
- CUPS
- Rocket.Chat
- PHP / Magento
How I work
Four steps, always in this order
No audit means no design, no design means no rollout. This order saves money by the third step.
Step 01
Audit and analysis
An inventory of what is actually running and a risk list ordered by business impact. Without this step everything that follows is guesswork.
Step 02
Architecture design
The target shape of the environment, with cost, schedule and a rollback plan. Decisions are made before the rollout, not during it.
Step 03
Implementation
Staged delivery, with configuration described in Ansible and Terraform. Every step is repeatable and reversible.
Step 04
Operations and monitoring
Monitoring with thresholds tuned to real traffic, updates, restore testing and a report on the state of the environment.
Frequently asked questions
Contact
Describe the problem and get a specific answer
The fastest way to get to the point is to outline your current environment and what needs to change in your first message.