SRE Engineer
Indexed description
Key Responsibilities
Infrastructure & Platform Engineering
- Ensure system availability and reliability through automated monitoring strategies
- Produce post-mortems and implement resulting process improvements
- Proactively mitigate operational risks through risk assessment and wider collaboration with
- engineering teams
- Design, implement, and continuously measure and improve risk mitigation strategies
- Monitor system health through observability and telemetry
- Unblock bottlenecks in system performance
- Minimize emergency response time priods
- Maintain internal tooling surrounding bug tracking, CI/CD pipelines, and wider
- communication with the teams
- Partner with Dev, DevOps, and QA teams to resolve infrastructure or deployment blockers during release cycles
- Provide technical guidance and mentorship to platform engineers
- Participate in architectural reviews, release readiness checkpoints, and root-cause analyses
- 10+ years of experience in infrastructure, platform, or DevOps engineering roles
- Strong programming skills (e.g., Python, Go)
- Hands-on experience with hybrid infrastructure (cloud + bare metal)
- Deep knowledge of infrastructure-as-code tools (Terraform, Helm, Kubernetes)
- Proven cross-functional collaboration skillsPreferred Qualifications
- Experience with GPU-accelerated compute or HPC-style infrastructure
- Familiarity with platform engineering or developer experience optimization
- Experience with high-velocity release cycles and incident response
- Experience with GPU-accelerated compute or HPC-style infrastructure
- Familiarity with platform engineering or developer experience optimization
- Experience with high-velocity release cycles and incident response
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search