Back to search
Quadrant IQ Solutions LLC Linkedin · Posted today

AWS Administrator

Columbus

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

AWS & Site Reliability Engineering

  • Design, deploy, and maintain highly available, scalable, and fault-tolerant applications and infrastructure on AWS.
  • Implement SRE principles including SLIs, SLOs, SLAs, error budgets, availability, reliability, and capacity planning.
  • Manage AWS services including EC2, EKS, ECS, Lambda, S3, RDS, DynamoDB, API Gateway, CloudFront, Route 53, IAM, VPC, CloudWatch, SNS, SQS, and EventBridge.
  • Troubleshoot complex production issues involving application, infrastructure, networking, database, and cloud components.
  • Participate in incident management, root-cause analysis, problem management, and post-incident reviews.
  • Develop automation to reduce manual operational activities and eliminate repetitive tasks.
  • Perform capacity planning, performance tuning, disaster recovery, and business continuity activities.

DevOps & Infrastructure Automation

  • Build and maintain CI/CD pipelines using tools such as Jenkins, GitHub Actions, GitLab CI/CD, AWS CodePipeline, or Azure DevOps.
  • Implement Infrastructure as Code using Terraform, CloudFormation, or AWS CDK.
  • Automate infrastructure provisioning, configuration management, deployments, and operational processes.
  • Implement containerized workloads using Docker and Kubernetes/Amazon EKS.
  • Develop automated deployment strategies including blue-green, canary, and rolling deployments.
  • Integrate security, compliance, testing, and quality checks into CI/CD pipelines.

AIOps & AI-Driven Operations

  • Implement AIOps solutions to improve monitoring, incident detection, event correlation, root-cause analysis, and automated remediation.
  • Leverage AI/ML and Generative AI capabilities to analyze logs, metrics, traces, alerts, and operational data.
  • Develop intelligent alerting and anomaly-detection mechanisms to identify potential production issues before they impact customers.
  • Build AI-assisted incident investigation and troubleshooting workflows.
  • Integrate LLM/GenAI capabilities into SRE and DevOps workflows for automated log analysis, incident summarization, knowledge retrieval, and remediation recommendations.
  • Develop automated runbooks and self-healing mechanisms using event-driven AWS services and AI-assisted decision making.
  • Integrate AIOps platforms and observability tools such as Datadog, Dynatrace, New Relic, Splunk, CloudWatch, Grafana, and Prometheus.
  • Develop or integrate AI agents/workflows that can assist with incident response, operational diagnostics, and infrastructure management.
  • Monitor AIOps/AI solutions for accuracy, reliability, security, and operational effectiveness.

Observability & Monitoring

  • Implement comprehensive metrics, logs, traces, dashboards, and alerting across cloud and application environments.
  • Work with Prometheus, Grafana, CloudWatch, OpenTelemetry, Datadog, Dynatrace, Splunk, or similar observability platforms.
  • Establish meaningful service-level indicators and operational dashboards.
  • Implement distributed tracing and application performance monitoring.
  • Tune alerts to reduce false positives and alert fatigue.
  • Build proactive monitoring and predictive health checks.


Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search