DevOps - Site Reliability Engineer
Indexed description
Essential Functions
- Pipeline Engineering:
deployment processes for both web-based IoT applications and firmware.
Collaborate with engineering teams to implement effective branching strategies, code
reviews, and automated testing frameworks.
Continuously optimize pipeline performance and reliability to accelerate delivery cycles.
- Cloud Infrastructure Engineering:
scaling of resources.
Automate routine tasks and implement infrastructure as code (IaC) practices to improve
efficiency and consistency.
Monitor system performance and proactively identify and resolve potential issues.
Implement robust security measures to protect our cloud environment.
- Site Reliability Engineering
Respond to incidents and outages promptly, implementing effective incident response
procedures.
Analyze system logs and metrics to identify potential issues and bottlenecks.
Collaborate with development teams to improve software quality and reduce failure rates.
- Strong proficiency in scripting languages (Python, Bash, etc.) and configuration management
- Experience with CI/CD tools (Jenkins, GitHub CI/CD, GitHub Workflows and Actions, etc.) and
- Deep understanding of cloud platforms, particularly AWS
- Knowledge of containerization technologies (Docker, Kubernetes) and orchestration tools
- Knowledge of containerization security best practices, tools and techniques.
- Knowledge of storage administration – NFS/EFS.
- Knowledge of disaster recovery best practices and tools.
- Knowledge of network security best practices and tools.
- Experience with infrastructure change management best practices.
- Experience with infrastructure as code (IaC) practices (Terraform, CloudFormation)
- Solid understanding of networking concepts (TCP/IP, DNS, load balancing)
- Familiarity with monitoring and logging tools (Prometheus, Grafana, ELK Stack)
- Strong problem-solving and troubleshooting skills
- A passion for automation and continuous improvement
Education: BS degree in Computer Science, Software Engineering or relevant field – or equivalent experience.
Experience
Minimum 5 years total experience in progressively responsible positions in the field of specialty.
Excellent analytical and time management skills
Teamwork skills with a problem-solving attitude
Experience working with other engineers to define constraints and requirements.
Professional And Organizational Skills Are Essential.
Ability to effectively communicate at all levels of the organization and externally (ie. customers, suppliers, auditors, etc.) through good written and verbal communication skills.
Ability to multitask and manage a variety of assigned projects.
Team building skills and the ability to foster an environment of cooperation and teamwork.
Ability to effectively and appropriately delegate workload.
Ability to prioritize tasks, direct the tasks of subordinates, and achieve objectives with minimal supervision.
Propensity to listen and approach new ideas and challenges with an open mind and evaluate situations on the merit of fact and data.
Summary
The DevOps/Site Reliability Engineer will lead and drive all aspects of the infrastructure supporting our Internet of Things (IoT) application servicing controlled dispensing solutions across the world. This includes operational environments – production, staging, testing, development and others as needed, the continuous integration/continuous deliver pipeline, securing the production environment, and optimization of the production environment.
The DevOps/Site Reliability Engineer will play a pivotal role in bridging the gap between software development and IT operations. They are responsible for ensuring the reliability, security, performance, and scalability of complex systems. Their focus lies in automating processes, implementing continuous integration and delivery pipelines, and proactively addressing potential issues.
A key aspect of their work is fostering collaboration between development and operations teams. They promote a culture of shared responsibility, breaking down silos and enabling faster and more efficient software delivery. By leveraging automation and monitoring tools, the DevOps/Site Reliability Engineer streamlines workflows, reduce manual intervention, and minimize downtime.
The DevOps/Site Reliability Engineer is the architects of reliable and resilient systems. They willo continuously strive to improve system performance, enhance security, and ensure seamless user experiences.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search