Senior Site Reliability Engineer
Indexed description
Description
• Operate and manage large-scale systems with high availability and resilience requirements.
• Build internal tools and scripts to eliminate manual work and SRE/DevOps tasks.
• Automate infrastructure provisioning, configuration, deployment, and monitoring across on-premise and cloud (e.g., AWS) environments.
• Collaborate with development teams to design and maintain scalable, reliable, and secure systems.
• Apply security and compliance best practices (e.g., PCI DSS, ISO 27001) across infrastructure.
• Monitor and respond to incidents 24/7 with a focus on root cause elimination.
• Continuously improve system performance, scalability, and reliability.
Requirement
• 5+ years of experience in SRE, DevOps, or Infrastructure Engineering roles.
• Strong Linux systems background with solid understanding of OS-level debugging and performance tuning.
• Expertise in CI/CD and automation tools (e.g., Jenkins, GitLab CI, Terraform, ArgoCD, Prometheus, Grafana).
• Experience or strong interest in integrating AI Agents into SRE workflows for system monitoring, log analysis, and incident response automation.
• Proficient in scripting languages such as Python, Go, or Bash.
• Deep knowledge of Kubernetes, container orchestration, and containerization best practices.
• Familiarity with microservices architecture, observability, service mesh, and API gateways.
• Experience with distributed systems technologies such as Kafka, Redis, MySQL, MongoDB, ETCD.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search