Platform Reliability Engineer
Indexed description
Role name : Platform Reliability Engineer (Full Stack & Cloud Native)
Skills required (Key skill : Zabbix)
Expertise in core tool sets including (but not limited to): Go, Python 3, FastAPI, React, HTML, JSON, JavaScript, NGINX, Docker, Podman, Kafka, Grafana, Air Flow, Linux, basic networking, Databases (MySQL, ClickHouse, Victoria Metrics, Postgres, MongoDB & OpenTSDB), and experience in other open source technologies like Telemetry, Zabbix, Kubernetes container and NOC.
The current monitoring platform, NOC, is an open source network management system providing discovery, inventory, fault management, and SNMP based performance monitoring capabilities. To simplify operations, standardise monitoring, and improve observability, SNMP monitoring functions need to be migrated to Zabbix while decommissioning or reducing reliance on NOC for monitoring functions.
Objectives
- Replace NOC SNMP performance monitoring with Zabbix.
- Ensure continuous visibility of network devices via SNMP polling.
- Standardise monitoring templates and alerting models.
- Improve operational efficiency for NOC teams.
- Reduce operational complexity by consolidating monitoring tools.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search