AWS Administrator
Indexed description
AWS & Site Reliability Engineering
- Design, deploy, and maintain highly available, scalable, and fault-tolerant applications and infrastructure on AWS.
- Implement SRE principles including SLIs, SLOs, SLAs, error budgets, availability, reliability, and capacity planning.
- Manage AWS services including EC2, EKS, ECS, Lambda, S3, RDS, DynamoDB, API Gateway, CloudFront, Route 53, IAM, VPC, CloudWatch, SNS, SQS, and EventBridge.
- Troubleshoot complex production issues involving application, infrastructure, networking, database, and cloud components.
- Participate in incident management, root-cause analysis, problem management, and post-incident reviews.
- Develop automation to reduce manual operational activities and eliminate repetitive tasks.
- Perform capacity planning, performance tuning, disaster recovery, and business continuity activities.
DevOps & Infrastructure Automation
- Build and maintain CI/CD pipelines using tools such as Jenkins, GitHub Actions, GitLab CI/CD, AWS CodePipeline, or Azure DevOps.
- Implement Infrastructure as Code using Terraform, CloudFormation, or AWS CDK.
- Automate infrastructure provisioning, configuration management, deployments, and operational processes.
- Implement containerized workloads using Docker and Kubernetes/Amazon EKS.
- Develop automated deployment strategies including blue-green, canary, and rolling deployments.
- Integrate security, compliance, testing, and quality checks into CI/CD pipelines.
AIOps & AI-Driven Operations
- Implement AIOps solutions to improve monitoring, incident detection, event correlation, root-cause analysis, and automated remediation.
- Leverage AI/ML and Generative AI capabilities to analyze logs, metrics, traces, alerts, and operational data.
- Develop intelligent alerting and anomaly-detection mechanisms to identify potential production issues before they impact customers.
- Build AI-assisted incident investigation and troubleshooting workflows.
- Integrate LLM/GenAI capabilities into SRE and DevOps workflows for automated log analysis, incident summarization, knowledge retrieval, and remediation recommendations.
- Develop automated runbooks and self-healing mechanisms using event-driven AWS services and AI-assisted decision making.
- Integrate AIOps platforms and observability tools such as Datadog, Dynatrace, New Relic, Splunk, CloudWatch, Grafana, and Prometheus.
- Develop or integrate AI agents/workflows that can assist with incident response, operational diagnostics, and infrastructure management.
- Monitor AIOps/AI solutions for accuracy, reliability, security, and operational effectiveness.
Observability & Monitoring
- Implement comprehensive metrics, logs, traces, dashboards, and alerting across cloud and application environments.
- Work with Prometheus, Grafana, CloudWatch, OpenTelemetry, Datadog, Dynatrace, Splunk, or similar observability platforms.
- Establish meaningful service-level indicators and operational dashboards.
- Implement distributed tracing and application performance monitoring.
- Tune alerts to reduce false positives and alert fatigue.
- Build proactive monitoring and predictive health checks.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search