Back to search
CLOUDSUFI Linkedin · Posted 4d ago

Junior SRE Engineer

Guadalajara

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Junior SRE Engineer – AI-Driven SRE & Cloud SRE

Location: Country Club, Guadalajara, Jalisco, Mexico 44610

Experience: 1–3 Years


Education: BTech / BE / MCA / MSc Computer Science

Reporting To: Lead SRE / Solution Architect – Reliability Engineering

About the Role

CLOUDSUFI is looking for a Junior SRE Engineer to join an AI-driven Site Reliability Engineering team supporting a regulated enterprise platform in Guadalajara.

This role is ideal for an early-career SRE, DevOps, Cloud, or Production Support Engineer who wants hands-on exposure to AWS, Datadog, Terraform, Kubernetes, Python/Bash automation, CI/CD, incident management, and AI-driven reliability engineering.

You will work closely with senior SRE engineers and architects to support production reliability, observability, automation, cloud operations, and AI-enabled incident response.

Key Responsibilities

1. AI-Augmented Incident Response & RCA

  • Support incident response as a secondary/shadow responder under senior SRE guidance.
  • Analyze metrics, logs, traces, and deployment history to help identify incident causes.
  • Support AI-driven incident detection, correlation, summarization, and RCA tools.
  • Assist with postmortems and tracking corrective and preventive actions.

2. Observability & Alert Management

  • Build and maintain Datadog monitors, dashboards, and SLOs using Terraform.
  • Support anomaly detection, outlier detection, forecast monitors, composite alerts, and dynamic thresholds.
  • Help reduce alert noise and improve monitor quality.
  • Conduct monitor-hygiene and coverage-gap reviews.
  • Support SLI/SLO and error-budget reviews for critical services.

3. AWS Cloud Reliability

  • Support AWS workloads including ECS/Fargate, EKS, Lambda, RDS/Aurora, ALB, SQS/SNS, and Step Functions.
  • Monitor capacity, performance, saturation, and cloud-cost trends.
  • Assist with cloud observability and cost governance.
  • Identify unusual or high-cost telemetry and infrastructure patterns and escalate findings.

4. AI Governance & Reliability

  • Participate in reliability and architecture reviews under senior guidance.
  • Learn how risks such as prompt injection, authorization gaps, secrets exposure, and autonomous-agent blast radius are assessed.
  • Support production-readiness, change-management, and audit requirements.
  • Assist in tracking AI-risk and reliability remediation activities.

5. DevOps, CI/CD & Automation

  • Support reliability and security controls within CI/CD pipelines.
  • Assist with AI-enabled change-risk, configuration, dependency, and infrastructure-drift checks.
  • Develop automation using Python and Bash against AWS, Datadog, GitHub, PagerDuty, and Jira APIs.
  • Contribute to Terraform modules and pull requests.
  • Help create runbooks and progressively automate operational procedures.

6. Reliability Workflow Support

  • Support multiple reliability initiatives across SRE and Security teams.
  • Maintain operational documentation and workflow tracking.
  • Contribute to a reliability-first, automation-focused engineering culture.

Required / Relevant Technical Skills

  • 1–3 years of experience in SRE, DevOps, Cloud, Platform Engineering, or Production Support.
  • Exposure to AWS or another major cloud platform.
  • Knowledge of monitoring/observability concepts: metrics, logs, traces, and APM.
  • Exposure to Datadog, Grafana, Prometheus, New Relic, CloudWatch, ELK, or OpenSearch.
  • Basic understanding of SLI, SLO, error budgets, and alert management.
  • Knowledge of Docker and Kubernetes/ECS.
  • Beginner-to-intermediate Terraform, Ansible, or CloudFormation experience.
  • Basic Linux and networking troubleshooting skills.
  • Understanding of DNS, TLS, load balancing, timeouts, and retries.
  • Some hands-on experience with Python, Bash, or similar scripting.
  • Understanding of REST APIs and JSON.
  • Experience with Git and pull-request workflows.
  • Exposure to CI/CD tools such as GitHub Actions, Jenkins, GitLab CI, or ArgoCD.
  • Familiarity with incident management, escalation processes, and PagerDuty/Opsgenie.
  • Interest in AI/LLM-based operational tools and AIOps.

Core Competencies

  • SRE fundamentals: SLI, SLO, error budgets, and toil reduction
  • Observability and monitoring
  • AIOps and anomaly detection
  • Incident response and RCA
  • AWS cloud infrastructure
  • Docker, ECS, and Kubernetes
  • Terraform / Infrastructure as Code
  • CI/CD and release engineering
  • Python/Bash automation
  • Cloud cost and reliability engineering
  • Resilience patterns
  • Disaster recovery and failure-management concepts
  • Runbook automation and self-healing
  • DevSecOps fundamentals
  • AI-agent risk and governance fundamentals

Preferred Certifications

Certifications are preferred but not mandatory:

  • AWS Certified Cloud Practitioner or Associate
  • HashiCorp Terraform Associate
  • Datadog Fundamentals or equivalent
  • CKA or KCNA
  • AI/ML or AIOps-related certification/exposure

Ideal Candidate

We are looking for someone who:

  • Has 1–3 years of hands-on technical experience.
  • Enjoys troubleshooting problems using data and evidence.
  • Is interested in cloud infrastructure and reliability engineering.
  • Wants to grow in AWS, Terraform, Datadog, Kubernetes, and automation.
  • Is curious about applying AI to SRE and operational workflows.
  • Understands the importance of production discipline and incident escalation.
  • Is comfortable working in a US shift and participating in a shadow on-call rotation.
  • Communicates clearly and works well with senior engineers and architects.

Location: Country Club, Guadalajara, Jalisco, Mexico 44610

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search