QA Automation Engineer - Cloud & Infrastructure
Indexed description
You will build automation that validates complex cloud-native systems across Kubernetes and multi-cloud environments. You'll test not just whether a feature works, but whether the underlying infrastructure remains reliable under scale, failure, and operational stress.
What You’ll Work On
- Kubernetes-based distributed systems
- AWS, Azure and GCP environments
- Multi-cloud and multi-cluster infrastructure
- Infrastructure and deployment automation
- Kubernetes workload and cluster validation
- Cloud service integrations
- Observability and alerting pipelines
- Reliability validation and chaos testing
- API, integration and end-to-end automation
- CI/CD reliability gates
- Design and build scalable automation frameworks for cloud infrastructure and system-level validation
- Automate validation of Kubernetes clusters, workloads and deployments
- Test cloud infrastructure across AWS, Azure and GCP
- Validate infrastructure changes and deployment workflows
- Build automated tests for multi-cluster and distributed environments
- Design failure-injection and chaos scenarios for cloud infrastructure
- Validate system behavior during infrastructure and service failures
- Integrate automated validation into CI/CD pipelines
- Analyze failures using logs, metrics and traces
- Partner with SRE, Platform and Backend teams to improve system reliability and testability
- Lead root cause analysis of infrastructure and production failures
- Build reusable automation utilities and validation tooling
- Establish reliability gates that prevent defective infrastructure or deployments from reaching production
- Cloud infrastructure and distributed systems
- Kubernetes architecture and troubleshooting
- Cloud failure modes and reliability
- Multi-cluster environments
- Observability and production debugging
- Infrastructure automation
- AWS / Azure / GCP
- EKS, GKE or managed Kubernetes platforms
- Docker and Kubernetes
- Infrastructure-as-Code
- CI/CD pipelines
- Cloud networking fundamentals
- Chaos testing and reliability engineering
- API and system-level testing
- Strong coding skills in Python and Go (mandatory)
- Experience building automation frameworks and system-level tooling
- Proficiency in Shell scripting and infrastructure automation
- You won't just test whether an API returns the expected response. You'll be testing how a distributed cloud-native system behaves under real-world conditions — at scale, across clusters, and when things fail.
- You'll build automation to answer questions like:
- What happens when infrastructure fails?
- Can the system detect it?
- Does it recover correctly?
- Can we validate that behavior automatically before it reaches production?
- The goal isn't just to find bugs. It's to build systems that can prove they are reliable.
Our team includes experienced entrepreneurs and engineers who have built multiple billion-dollar products from scratch. As a well-funded US-based company backed by top-tier VCs, we have offices in the US, India, and Europe. Join us in our fast-paced environment where you’ll have a front-row seat to shape the future of AI-driven Observability solutions.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search