Network Observability/Monitoring Engineer
Indexed description
We are seeking a Network Monitoring Platform Engineer to support and enhance our large telecommunication client's customized network monitoring and service assurance platform. This platform collects device health and performance data for end customers and has been heavily customized to support internal specific policies, protocols, and operational standards.
This role will focus on platform enhancements, upgrades, service migrations, troubleshooting, and ongoing operational support. The ideal candidate will have experience supporting monitoring or observability platforms, strong Linux administration skills, and the ability to develop, enhance, and deploy services in production environments.
Core Responsibilities
- Enhance, tune, and maintain an enterprise network monitoring and service assurance platform
- Develop and maintain Python-based services and integrations
- Troubleshoot production issues and provide operational support as needed
- Configure, deploy, upgrade, and manage platform instances across environments
- Build and maintain REST API integrations for custom platform functionality
- Support monitoring and observability capabilities using Grafana and Prometheus
- Develop automation and deployment processes using Ansible, Bash, Jenkins, and CI/CD pipelines
- Perform platform upgrades, migrations, and service modernization efforts
- Troubleshoot platform performance, data collection, and service health issues
- Support Linux-based infrastructure running on Red Hat Enterprise Linux
- Work with relational databases to support platform services and integrations
- Collaborate with engineering teams to implement new monitoring standards and capabilities
- Configure and maintain monitoring services that capture network and device health information from customer environments
Required Qualifications
- 5+ years supporting network monitoring, observability, service assurance, NOC, or telecommunications platforms
- Strong Python development experience (required)
- Experience building and consuming REST APIs
- Strong Linux administration experience, preferably Red Hat Enterprise Linux
- Experience with Prometheus and Grafana
- Experience with Ansible, Bash scripting, and CI/CD pipelines using Jenkins
- Experience troubleshooting and supporting applications in production environments
- Experience deploying, configuring, upgrading, and maintaining platform services
- Understanding of networking fundamentals and how device health and telemetry data are collected and monitored
- Experience working with relational databases
Preferred Qualifications
- Experience with Oracle Unified Assurance (Assure1), Netcool, SevOne, ScienceLogic, SolarWinds, LogicMonitor, Splunk, Dynatrace, or similar monitoring platforms
- Telecommunications, NOC, or OSS/BSS experience
- Java development experience
- Experience with fault management, event correlation, and service assurance concepts
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search