Zabbix to Grafana Migration

If your team has stopped reading the alerts, the problem is not the tool. It is the alarm model.

Teams that have run Zabbix for a year or more rarely have a monitoring problem. They have an alarm problem: the platform is collecting correctly, but it produces so many notifications that operators have learned to ignore the channel they arrive on. At that point the monitoring system has stopped doing its job even though every check is green.

GİTA Teknoloji migrates these environments to a Grafana-based alerting and reporting model. In most cases Zabbix stays exactly where it is as the collection layer, and what changes is how alarms are defined, grouped, suppressed and delivered. That distinction matters commercially: it is a far smaller project than replacing a monitoring platform.

The deliverable is not a dashboard. It is a smaller number of alarms that are each worth acting on, plus weekly and monthly reports that turn service quality into a figure your management can read.

Signs a migration is warranted

If two or more of these are familiar, the alarm model is the bottleneck.

Notification volume nobody reads

Hundreds or thousands of notifications a day. Operators mute the channel, and a real outage arrives in the same stream as the noise.

One failure, many alarms

A single upstream link drop pages for every device behind it, because dependencies were never modelled.

Flapping links treated as incidents

Marginal WAN circuits generate up-down-up cycles all day, each one a separate alarm.

No reporting layer

You can see the current state, but you cannot answer "how available was site 214 last month" without manual work.

Dashboards nobody opens

Default templates that show everything, which in practice means they show nothing anyone is looking for.

Alarms with no owner

Notifications go to a shared inbox or group chat with no escalation path, so responsibility is ambiguous by design.

How we run the migration

Staged, with a parallel run, so there is no window where the estate is unmonitored.

1. Alarm audit

We export what currently fires, over a representative period, and classify it: acted on, ignored, duplicate, or symptom of another alarm. This is usually the moment the scale of the noise becomes visible.

2. Decide what stays

Zabbix usually remains as the collector. Where the estate is container- or cloud-heavy we add Prometheus alongside it rather than replacing anything.

3. Rebuild the alarm model

Thresholds set against observed behaviour rather than defaults, dependency chains so an upstream failure suppresses downstream noise, maintenance windows, and flap damping on known-marginal links.

4. Dashboards by audience

A NOC view for the people on shift, a service view for management, and per-site views for the teams who own individual locations.

5. Parallel run

Old and new alerting run side by side until the new model has been proven against real incidents. Nothing is switched off on the strength of a demo.

6. Cutover and handover

Notification channels move across, runbooks are written for the alarms that remain, and the platform either goes to your team or stays with ours under a managed service.

What changes for the team

Fewer, actionable alarms

The measure of success is not how much you can see. It is how few notifications a shift receives, and how many of those need someone to do something.

Alarms that arrive where people are

Delivery to the channel the team already uses — Telegram, email, SMS for critical events, or a chat platform — rather than a channel they have to remember to check.

Reporting without manual work

Weekly and monthly availability and incident reports generated on a schedule and delivered by email.

Dependency-aware escalation

A branch router going down produces one alarm naming the branch, not one alarm per device behind it.

History that survives the move

Keeping the existing collector in place means existing history stays queryable rather than being abandoned at cutover.

A documented model

Alarm definitions and dependencies held as code, so the reasoning is reviewable rather than living in one engineer’s memory.

Frequently asked questions

Do we have to abandon Zabbix?

Usually not. Zabbix is a capable collector, and in most migrations it stays. What we replace is the alerting and visualisation layer, which is where the pain actually is. Replacing the collector as well is a decision that comes out of discovery, not one we assume.

Will we lose our historical data?

No, when the existing collector remains in place its history remains queryable. Where a collector genuinely has to be replaced, carrying history across is scoped explicitly as part of the deployment rather than assumed either way.

Can we keep using Telegram for alarms?

Yes. Teams already using Telegram normally keep it. Changing the notification channel and the alarm model at the same time makes it impossible to tell which change produced the improvement.

How long does a migration take?

It depends on estate size and how much of the alarm audit turns up. A single-site estate is a matter of weeks; a several-hundred-site WAN estate is longer, and the parallel-run period is the part that should not be compressed.

Can you keep running it afterwards?

Yes. Most customers take the migration and the ongoing managed service together, with a defined support window, ticket allowance and escalation path. Handing the platform to your own team is equally supported.

What if our routers do not expose SNMP?

ICMP keepalives and TCP probes give availability and latency without SNMP. Where SNMP is available we add interface-level metrics, capacity and hardware health, which is a meaningfully richer picture — so establishing what the devices expose is part of discovery.

Start with the alarm audit

Send us a week of your current notification volume and we will tell you how much of it is signal.

Get in touch