Principal Site Reliability Engineer
Indexed description
Job Family Definition:
Designs, develops, troubleshoots and debugs software programs for software enhancements and new products. Develops software including operating systems, compilers, routers, networks, utilities, databases and Internet-related tools. Determines hardware compatibility and/or influences hardware design.
Management Level Definition:
Contributions impact technical components of HPE products, solutions, or services regularly and sustainable. Applies advanced subject matter knowledge to solve complex business issues and is regarded as a subject matter expert. Provides expertise and partnership to functional and technical project teams and may participate in cross-functional initiatives. Exercises significant independent judgment to determine best method for achieving objectives. May provide team leadership and mentoring to others.
In a typical day as a Principal Site Reliability Engineer, you would...
As a Principal Site Reliability Engineer, you will play a key role in designing, building, and optimizing cloud infrastructure and deployment systems. Your work will directly impact scalability, security, and operational efficiency across our platforms. Key responsibilities include:
- Enhance Infrastructure as Code (IAC) and enforce best practices.
- Optimize cloud infrastructure for scalability, security, and cost-effectiveness.
- Develop internal tools to support and streamline cloud platform operations.
- Improve CI/CD pipelines and deployment workflows using FluxCD and Jenkins.
- Address container image vulnerabilities and standardize remediation processes.
- Build Amazon Machine Images (AMIs) aligned with CIS and STIG benchmarks.
- Strengthen monitoring, alerting, and observability using Prometheus, Grafana, and logging tools.
- Troubleshoot complex production issues to ensure system reliability and customer satisfaction.
- Fine-tune distributed systems such as Apache Kafka and Cassandra.
- Collaborate with development, security, and operations teams to align infrastructure with application needs.
What you need to bring:
- Minimum of 10 years of hands-on experience in Infra Ops, Dev Ops, or Site Reliability Engineering (SRE).
- Proficiency with Linux systems, especially Debian-based distributions.
- Strong experience with cloud platforms such as AWS and GCP.
- Expertise in Infrastructure as Code tools like Terraform, Packer, and Ansible.
- Solid programming skills in Python and/or Golang.
- Deep understanding of containerization (Docker, Container) and orchestration tools (AWS EKS, GCP GKE).
- Experience with GitOps workflows.
- Proven track record in implementing and maintaining CI/CD pipelines.
- Strong background in security and familiarity with security programs.
- Experience with monitoring and logging tools (Prometheus, Grafana, ELK).
- Knowledge of both relational (SQL) and non-relational databases.
- Excellent problem-solving and debugging skills with a strong sense of ownership.
- Experience managing distributed systems like Apache Kafka and Cassandra.
- Effective communicator and collaborative team player.
- It is mandatory to attend to San Juan office twice a week.
Preferred Qualifications
- Experience contributing to open-source projects.
- Background in security engineering or related disciplines.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search