AWS SRE – AIOps & Full Stack Engineer
Indexed description
The ideal candidate will have a strong background in AWS infrastructure, Site Reliability Engineering, DevOps, observability, AIOps, automation, and full-stack application development. This role requires someone who can work across the entire technology stack—from cloud infrastructure and CI/CD pipelines to backend services, APIs, frontend applications, monitoring, and AI-driven operational automation.
The engineer will be responsible for improving system reliability, scalability, performance, and operational efficiency while leveraging AI/ML and AIOps capabilities to proactively identify incidents, predict failures, automate remediation, and improve application observability.
Key Responsibilities
AWS & Site Reliability Engineering
- Design, deploy, and maintain highly available, scalable, and fault-tolerant applications and infrastructure on AWS.
- Implement SRE principles including SLIs, SLOs, SLAs, error budgets, availability, reliability, and capacity planning.
- Manage AWS services including EC2, EKS, ECS, Lambda, S3, RDS, DynamoDB, API Gateway, CloudFront, Route 53, IAM, VPC, CloudWatch, SNS, SQS, and EventBridge.
- Troubleshoot complex production issues involving application, infrastructure, networking, database, and cloud components.
- Participate in incident management, root-cause analysis, problem management, and post-incident reviews.
- Develop automation to reduce manual operational activities and eliminate repetitive tasks.
- Perform capacity planning, performance tuning, disaster recovery, and business continuity activities.
- Build and maintain CI/CD pipelines using tools such as Jenkins, GitHub Actions, GitLab CI/CD, AWS CodePipeline, or Azure DevOps.
- Implement Infrastructure as Code using Terraform, CloudFormation, or AWS CDK.
- Automate infrastructure provisioning, configuration management, deployments, and operational processes.
- Implement containerized workloads using Docker and Kubernetes/Amazon EKS.
- Develop automated deployment strategies including blue-green, canary, and rolling deployments.
- Integrate security, compliance, testing, and quality checks into CI/CD pipelines.
- Implement AIOps solutions to improve monitoring, incident detection, event correlation, root-cause analysis, and automated remediation.
- Leverage AI/ML and Generative AI capabilities to analyze logs, metrics, traces, alerts, and operational data.
- Develop intelligent alerting and anomaly-detection mechanisms to identify potential production issues before they impact customers.
- Build AI-assisted incident investigation and troubleshooting workflows.
- Integrate LLM/GenAI capabilities into SRE and DevOps workflows for automated log analysis, incident summarization, knowledge retrieval, and remediation recommendations.
- Develop automated runbooks and self-healing mechanisms using event-driven AWS services and AI-assisted decision making.
- Integrate AIOps platforms and observability tools such as Datadog, Dynatrace, New Relic, Splunk, CloudWatch, Grafana, and Prometheus.
- Develop or integrate AI agents/workflows that can assist with incident response, operational diagnostics, and infrastructure management.
- Monitor AIOps/AI solutions for accuracy, reliability, security, and operational effectiveness.
- Implement comprehensive metrics, logs, traces, dashboards, and alerting across cloud and application environments.
- Work with Prometheus, Grafana, CloudWatch, OpenTelemetry, Datadog, Dynatrace, Splunk, or similar observability platforms.
- Establish meaningful service-level indicators and operational dashboards.
- Implement distributed tracing and application performance monitoring.
- Tune alerts to reduce false positives and alert fatigue.
- Build proactive monitoring and predictive health checks.
- Develop and maintain scalable backend services, APIs, and web applications.
- Build RESTful APIs and microservices using technologies such as Java/Spring Boot, Python/FastAPI, Node.js, or similar.
- Develop responsive frontend applications using React, Angular, TypeScript, JavaScript, HTML, and CSS.
- Integrate frontend applications with REST/GraphQL APIs and cloud-native backend services.
- Design and optimize database interactions using PostgreSQL, MySQL, MongoDB, DynamoDB, or similar databases.
- Implement authentication and authorization using OAuth 2.0, OpenID Connect, JWT, AWS IAM, or similar technologies.
- Develop automated unit, integration, API, and end-to-end tests.
- Troubleshoot application performance and scalability issues across frontend, backend, database, and infrastructure layers.
- AWS
- EC2, S3, VPC, IAM, RDS, DynamoDB
- EKS/ECS, Lambda
- CloudWatch, Route 53, API Gateway
- SQS, SNS, EventBridge
- AWS networking and security
- Site Reliability Engineering principles
- Incident Management & Root Cause Analysis
- CI/CD
- Terraform / CloudFormation / AWS CDK
- Docker
- Kubernetes / EKS
- Jenkins / GitHub Actions / GitLab CI
- Linux administration and troubleshooting
- Bash/Shell scripting
- Python
- AIOps and intelligent event management
- Generative AI / LLM concepts
- AI-assisted incident management
- Anomaly detection and predictive monitoring
- Automated remediation / self-healing systems
- LLM APIs and AI agent workflows
- RAG/vector database concepts are a plus
- Experience integrating AI with DevOps/SRE workflows
- Datadog / Dynatrace / New Relic
- Prometheus
- Grafana
- Splunk
- AWS CloudWatch
- OpenTelemetry
- Application Performance Monitoring
- React / Angular
- JavaScript / TypeScript
- HTML / CSS
- Node.js / Python / Java
- REST APIs / GraphQL
- Microservices
- SQL / NoSQL databases
- Bachelor's degree in Computer Science, Engineering, Information Technology, or a related field.
- 7+ years of experience in software engineering, DevOps, Cloud Engineering, or SRE.
- 4+ years of hands-on AWS experience.
- Strong experience supporting production environments and distributed systems.
- Experience implementing AIOps or AI-powered operational solutions.
- Experience with Kubernetes and cloud-native architectures.
- Experience developing full-stack applications.
- Strong programming and scripting skills in Python, Java, Node.js, or similar languages.
- Experience working in Agile/Scrum environments.
- Strong troubleshooting, analytical, communication, and problem-solving skills.
- Cloud-native architecture
- Site Reliability Engineering
- AIOps & Generative AI
- Infrastructure automation
- Full-stack development
- Observability & monitoring
- Production support
- Incident response
- Automation & self-healing
- Performance optimization
- Security and reliability
- Continuous improvement
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search