SRE / Monitoring
Indexed description
Company Description Our organization is a technology-driven company focused on delivering reliable, scalable, and secure services to its customers. We strive to maintain high availability and performance across our platforms through modern infrastructure practices and continuous improvement. Teams collaborate closely across engineering, operations, and business functions to support innovation and stable growth. The environment emphasizes learning, professional development, and the adoption of best practices in observability, automation, and resilience.
Role Description This is a full-time, on-site SRE / Monitoring role based in Charlotte, NC. The role focuses on maintaining the reliability, performance, and availability of critical systems by designing and managing monitoring, alerting, and incident response processes. Day-to-day responsibilities include building and refining observability dashboards, analyzing system metrics and logs, identifying root causes of incidents, and implementing long-term fixes. The SRE will collaborate with software development and system administration teams to improve infrastructure, reduce manual work through automation, and ensure deployments meet reliability and scalability standards. The role also involves documenting operational runbooks, participating in on-call rotations, and contributing to continuous improvement of tools, processes, and reliability practices.
Qualifications
- Strong Site Reliability Engineering skills, including observability, incident management, and resilience-focused practices.
- Proficiency in Troubleshooting complex system issues across applications, infrastructure, and networks.
- Experience with Software Development, including scripting or coding for automation and tooling (e.g., Python, Go, or similar languages).
- Solid System Administration skills, with experience managing Linux or Unix-based environments and common services.
- Understanding of Infrastructure concepts, such as cloud platforms, containers, orchestration tools, and CI/CD pipelines.
- Familiarity with monitoring and logging tools (e.g., Prometheus, Grafana, Datadog, Splunk, or similar solutions).
- Ability to work collaboratively with cross-functional teams and communicate clearly in both technical and non-technical contexts.
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- Experience with security best practices, performance tuning, and capacity planning is beneficial.
- Previous involvement in on-call rotations and incident postmortem processes is a plus.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search