Introduction
High availability in IT operations requires proactive observability. A robust monitoring stack provides real-time alerting before network outages impact end-user experience.
In this technical breakdown, we outline the deployment of a centralized metrics pipeline using Prometheus, Node Exporter, SNMP Exporter, and Grafana.
Architecture & Data Flow
Below is the metric collection lifecycle from network nodes to Grafana dashboards:
[ Network Switches / Firewalls ] ---> (SNMP Exporter) --\
+---> (Prometheus TSDB) ---> (Grafana UI)
[ Linux / Proxmox Servers ] ---> (Node Exporter) --/ |
v
(Alertmanager)
Docker Compose Configuration
Create a docker-compose.yml file to orchestra the metrics engine:
version: '3.8'
services:
prometheus:
image: prom/prometheus:v2.45.0
container_name: prometheus
volumes:
- ./prometheus.yml:/etc/prometheus/prometheus.yml
- prom_data:/prometheus
ports:
- "9090:9090"
restart: unless-stopped
grafana:
image: grafana/grafana:10.0.0
container_name: grafana
ports:
- "3000:3000"
environment:
- GF_SECURITY_ADMIN_PASSWORD=SecuredDashboardPass123!
volumes:
- grafana_data:/var/lib/grafana
restart: unless-stopped
volumes:
prom_data:
grafana_data:
Prometheus Target Scraping
Configure prometheus.yml to define scrape intervals and targets:
global:
scrape_interval: 15s
evaluation_interval: 15s
scrape_configs:
- job_name: 'node_exporter'
static_configs:
- targets: ['10.0.30.15:9100', '10.0.30.16:9100']
- job_name: 'snmp_switches'
static_configs:
- targets: ['10.0.10.1', '10.0.10.2']
metrics_path: /snmp
params:
module: [if_mib]
Summary & Key Takeaways
- Scrape Interval Selection: Standardize on 15-second intervals for critical core devices and 60-second intervals for peripheral nodes.
- Persistence: Ensure persistent volumes are mounted for Prometheus TSDB data retention.