Campus Network Infrastructure
Designing and managing a scalable multi-site campus network infrastructure with VLAN segmentation, routing, and high availability.
- MikroTik
- Cisco
- VLAN
- OSPF
- BGP
- +2
An end-to-end telemetry and observability platform designed to monitor heterogeneous infrastructure including Linux servers, virtual machines, network routers, and storage arrays.
Centralized observability is crucial for proactive infrastructure management. This project unified disparate SNMP monitors, log files, and system metrics into single-pane-of-glass Grafana dashboards with automated Telegram and email notifications.
System admins previously relied on reactive ping checks and manual log inspection across individual servers, leading to delayed incident resolution and undetected disk space exhaustion outages.
Engineered a containerized monitoring stack leveraging Prometheus for time-series metrics, SNMP Exporter and LibreNMS for network equipment polling, Graylog for centralized syslog aggregation, and Grafana for visualization.
Data collectors poll SNMP devices and scrape Prometheus Node Exporters. Aggregated data is indexed and routed to Alertmanager and Grafana dashboards.
SNMP v3 polling for switch/router interfaces and Node Exporter / cAdvisor agents for host/container metrics.
Time-series database (Prometheus TSDB) and OpenSearch indexers for structured syslog events.
Custom Grafana dashboards grouped by service tier (Core Routers, VM Hosts, Storage Pools).
Prometheus Alertmanager rules with threshold triggers alerting on-call staff via webhook integration.
Deployed Docker Compose stack containing Prometheus, Grafana, Alertmanager, and Node Exporter.
Configured SNMP v3 credentials on core routers and switches for secure metric polling.
Authored custom Grafana dashboards for bandwidth utilization, CPU/RAM thresholds, and ping response latency.
Created alert rules for link downtime, packet loss spikes (>2%), CPU usage (>85% for 5 mins), and disk space capacity (<10%).
Integrated centralized rsyslog forwarding from Linux hosts to Graylog for log search and security auditing.
High SNMP polling frequency causing CPU load spikes on low-resource switches.
Adjusted polling intervals for non-critical SNMP OIDs from 15s to 60s while keeping interface bandwidth counters at 30s.
Excessive noisy alerts during scheduled backup windows.
Configured blackout maintenance silences in Alertmanager during automated backup cron schedules.
Designing and managing a scalable multi-site campus network infrastructure with VLAN segmentation, routing, and high availability.
Architecting a hardened Proxmox VE hypervisor cluster with Linux containerization, Nginx reverse proxy, and zero-trust VPN.
If you need assistance designing or auditing your network, hypervisor, or monitoring platform, let's talk.
Get In Touch