Site Reliability Engineer
Indexed description
Join our Site Reliability Engineering (SRE) team and work on large-scale infrastructure, automation, reliability, and cloud technologies.
Technology Stack
Linux: Red Hat, CentOS, Ubuntu
Containers & Cloud: Docker, OpenStack, oVirt
Automation & Configuration: Ansible, Puppet
Version Control: GitLab, Git
Networking & High Availability: BIND, Keepalived, HAProxy, Pacemaker, DRBD
Identity Management: FreeIPA
Monitoring & Logging: Prometheus, TICK Stack, Splunk, Nagios, Icinga
Other: Linux networking, infrastructure automation, scripting
Qualifications
- Bachelor's degree or higher in Computer Science, Computer Engineering, or a related technical field
- At least 5 years of experience supporting Linux/Unix systems
- Strong Linux administration and troubleshooting skills
- Good understanding of Linux networking and network internals
- Experience with Git, Ansible, Puppet, or similar automation/orchestration tools
- Understanding of Infrastructure as Code (IaC) concepts
- Experience with container technologies, particularly Docker
- Experience with Identity Management systems such as FreeIPA or Active Directory
- Experience with Bash, Python, Ruby, or other scripting/programming languages
- Ability to quickly learn and support different Unix/Linux environments
- Proactive mindset with the ability to identify problems, performance bottlenecks, and opportunities for improvement
- Strong documentation habits and a mindset of “document it once, solve it better next time”
- Willingness to provide occasional 24x7 support, including nights, weekends, and holidays, either onsite or remotely when required
- Business-level English required; business-level Japanese is a plus
- Experience operating OpenStack or other private cloud platforms is a plus
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search