Nipa Digital Marketing Agency
Linkedin · Posted yesterday
Site Reliability Engineer (SRE)
Continue to application
Add your email once, then Caio opens the original posting.
Indexed description
Responsibilities
- Collaborate with Product Owner and Tech Lead for planning technical activities.
- Collaborate with Development and Operations teams to improve CI/CD standards.
- Partner with development teams to ensure applications are designed with scale, resilience, and performance in mind.
- Participate in system design consulting, platform management production system through application reviews, testing, capacity planning, and performance tuning.
- Continuously seek out innovative ways to speed up and automate all aspects of testing, building, and releasing software project versions and updates through our infrastructure.
- Perform root cause analysis of production errors and resolve technical issues.
- Support team members throughout the cloud operations team.
- Act as a prime escalation point and Tier 2 cloud infrastructure engineer ticket (Freshdesk, Clickup).
- Propose automation ideas and solutions to reduce workload.
- Control internal portal change.
- Maintenance of SLA uptime.
- Study and take tests to acquire relevant certifications for SRE/DevOps.
- Other ad-hoc tasks as assigned by Manager.
- Bachelor’s degree in Computer Science, Computer Engineering or related technical major, or commensurate experience.
- 1-5 years in system engineer, systems development, SRE (Site Reliability Engineering), or Resilience Engineering.
- 1-5 years of server systems debug experience; debugging and root causing complex application platforms.
- Exposure to cloud or virtualization, Docker, Kubernetes (about AWS, Azure, Google Cloud, OpenStack, OpenShift, VMware vSphere) is plus.
- Experience with API, OS, Networking troubleshooting (API Request, CPU/Mem profiling, TCP/IP, DNS, routing, switching, firewalls, LAN/WAN or related).
- Interests in contributing towards increasing durability, security, availability and scalability of systems through exploration, diagnosis and remediation.
- Systematic thinking - ability to diagnose interactions between discrete components of system and drive product improvements.
- Familiarity with Linux (Ubuntu, Debian CentOS, RedHat) and Linux namespaces.
- Scripting in Ansible, Shell, Python3, Terraform, Jenkin to automate and operate system.
- Ability to handle and resolve critical customer issues/escalations, including on-call job rotations.
- Hands-on, self-starter and problem-solving attitude.
- Individual thinker and good team player.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search
Want help applying to roles like this?
Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search