Viper Search
Linkedin · Posted today
DevOps Developer
Continue to application
Add your email once, then Caio opens the original posting.
Indexed description
Site Reliability Engineer / DevOps
Role Overview
We are seeking a highly motivated Site Reliability Engineer (SRE) / DevOps professional to join
our team. As an interactive agency serving clients around the globe, our SRE team is critical to
maintaining high availability and performance worldwide. We operate on a "follow the sun"
model, leveraging our global team to ensure 24/7 coverage and rapid response.
The ideal candidate is proactive and thrives in a fast-paced, globally focused environment,
possessing a dual focus: rapid, effective incident response during critical events and proactive
development/automation when systems are operating normally.
Responsibilities
● Incident Response & Resolution: Act as a first responder to production incidents,
swiftly diagnosing, troubleshooting, and resolving issues to minimize downtime and
service disruption. Participate in a global, rotating on-call schedule as part of the "follow
the sun" team structure.
● Post-Incident Analysis: Conduct thorough root cause analyses (RCAs) of incidents,
implement permanent fixes, and integrate lessons learned into the development process
to prevent recurrence.
● System Automation: Design, implement, and maintain automation tools and
frameworks to streamline operational tasks, infrastructure provisioning, and continuous
deployment pipelines (CI/CD).
● Infrastructure Management: Manage and maintain our cloud infrastructure (AWS)
using Infrastructure as Code (IaC) tools (Terraform) to ensure consistency, reliability, and
security for global services.
● Monitoring and Alerting: Develop and refine comprehensive monitoring, logging, and
alerting strategies (e.g., using Datadog) to provide deep visibility into system health and
predict potential issues across all client regions.
● System Reliability & Performance: Collaborate with development teams to optimize
system performance and efficiency.
● Collaboration: Work closely with engineering teams across different time zones to
define reliability requirements, perform code reviews focusing on operational aspects,
and drive a culture of reliability and shared ownership.
Qualifications
● Proven experience as a DevOps Engineer, SRE, or similar role.
● Strong practical experience with AWS cloud services.
● Expertise in Infrastructure as Code (IaC), particularly Terraform.
Role Overview
We are seeking a highly motivated Site Reliability Engineer (SRE) / DevOps professional to join
our team. As an interactive agency serving clients around the globe, our SRE team is critical to
maintaining high availability and performance worldwide. We operate on a "follow the sun"
model, leveraging our global team to ensure 24/7 coverage and rapid response.
The ideal candidate is proactive and thrives in a fast-paced, globally focused environment,
possessing a dual focus: rapid, effective incident response during critical events and proactive
development/automation when systems are operating normally.
Responsibilities
● Incident Response & Resolution: Act as a first responder to production incidents,
swiftly diagnosing, troubleshooting, and resolving issues to minimize downtime and
service disruption. Participate in a global, rotating on-call schedule as part of the "follow
the sun" team structure.
● Post-Incident Analysis: Conduct thorough root cause analyses (RCAs) of incidents,
implement permanent fixes, and integrate lessons learned into the development process
to prevent recurrence.
● System Automation: Design, implement, and maintain automation tools and
frameworks to streamline operational tasks, infrastructure provisioning, and continuous
deployment pipelines (CI/CD).
● Infrastructure Management: Manage and maintain our cloud infrastructure (AWS)
using Infrastructure as Code (IaC) tools (Terraform) to ensure consistency, reliability, and
security for global services.
● Monitoring and Alerting: Develop and refine comprehensive monitoring, logging, and
alerting strategies (e.g., using Datadog) to provide deep visibility into system health and
predict potential issues across all client regions.
● System Reliability & Performance: Collaborate with development teams to optimize
system performance and efficiency.
● Collaboration: Work closely with engineering teams across different time zones to
define reliability requirements, perform code reviews focusing on operational aspects,
and drive a culture of reliability and shared ownership.
Qualifications
● Proven experience as a DevOps Engineer, SRE, or similar role.
● Strong practical experience with AWS cloud services.
● Expertise in Infrastructure as Code (IaC), particularly Terraform.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search
Want help applying to roles like this?
Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search