Lead DevOps Engineer
Indexed description
Job description
Title: Lead DevOps Engineer
Location: Wrexham/Hybrid/Remote
What we do:
We’re the leaders in outsourced calls, live chat and more, delivering brilliant conversations on behalf of businesses of all sizes. Fast-forward two decades and what started as a single, dedicated PA (who’s still with us today) looking after calls for a handful of local clients, is now a 1000-strong team working across two continents from our state-of-the-art UK headquarters in Wrexham, and our US office in Atlanta.
The role:
You will be a senior member of a small DevOps team managing and supporting multiple products hosted in AWS, Azure and GCP. As well as taking a hands-on approach to automating deployments of infrastructure and applications, performance monitoring, capacity management and disaster recovery, you will act as a technical lead for the team – setting standards, driving consistency in how we work, and mentoring other engineers.
We have a strategic objective to design and build highly resilient, cloud native services capable of dealing with increasing demand as the organisation grows in both the UK and US. You will be working closely with application development and infrastructure teams and will need to be able to translate business requirements into production ready solutions, while helping shape the technical direction of the platform.
Key responsibilities:
Technical Leadership & Standards
- Act as a team leader for the DevOps function, providing day-to-day technical direction and coordinating the team's work.
- Drive standardisation of tooling, patterns and processes across the team, reducing divergence and duplication.
- Own the documentation of best practices and engineering standards, and ensure they are followed consistently.
- Lead by example through high-quality, well-reviewed work, and raise the bar through constructive peer review.
- Mentor and support other engineers, helping with the adoption of new technologies and growing the team's capability.
- Evaluate new tools and approaches, and lead their adoption where they demonstrably improve how the team delivers.
Automation
- Lead the design and implementation of tooling to facilitate fast and efficient deployment and management of applications and infrastructure.
- Define and maintain the documented standards that operational processes, such as deployments and upgrades, must follow.
- Drive and prioritise an infrastructure as code approach to designing and implementing systems.
- Develop automated solutions for operational functions such as monitoring, performance and capacity management, and disaster recovery.
- Focus efforts on reducing toil and removing error by automating tasks and processes.
- Reinforce a culture of delivering quickly and effectively and iterating fast.
- Discover and document exceptions to turn manual work into repeatable actions and then into automation.
Observability
- Incorporate monitoring and logging features into systems during the design stage, ensuring that they become common components in each service.
- Build alerting systems that trigger on symptoms rather than on outages.
- Define and monitor Service Level Objectives (SLOs) for each service or product, and drive the team's adoption of SLO-based practices.
Security
- Ensure security best practices are followed during build and deployment of applications and infrastructure, and embed them into the team's standard patterns.
- Work with Information Security team members to ensure the hosting environment is secure.
- Lead the remediation of infrastructure security vulnerabilities detected during penetration testing and vulnerability scans, prioritising and coordinating work across the team.
Supporting the Business
- Debug and help resolve issues affecting the availability or performance of production systems, acting as a senior escalation point for complex problems.
- Participate in Post Incident Reviews to identify root cause and the actions required to prevent issues re-occurring, and ensure those actions are completed.
- Work to continuously improve the reliability, efficiency and scalability of systems and services.
- Continuously assess processes and develop ways to improve them.
- Build a deep, holistic understanding of business services and the underlying technology employed.
- Reduce MTTR by developing playbooks to resolve issues, and ensure runbook quality and coverage across the team.
- Work closely with application development teams to identify potential issues as early as possible within a product's lifecycle.
- Identify ways to reduce complexity and drive the standardisation of technology and processes where appropriate.
- Identify dependencies, bottlenecks and single points of failure, and work to mitigate risks.
- Participate in the planning of projects and work closely with infrastructure, operations, and application development teams to deliver objectives, providing input into technical design decisions.
- Participate in on-call support duties as part of the shared team rota.
The person:
- Significant hands-on production experience with at least two of the big three hyperscalers (AWS, Azure, GCP).
- Deep, demonstrable experience deploying, running and troubleshooting applications in Kubernetes in production environments.
- Strong track record of designing and operating GitOps-based delivery workflows and infrastructure as code at scale.
- Proficiency with CI/CD pipeline automation processes and tooling, including designing pipelines for other teams to consume.
- Extensive experience using declarative orchestration and configuration management tools in a production environment.
- Good understanding of common security issues, particularly the OWASP Top 10 CI/CD Security Risks, and experience embedding security controls into delivery pipelines.
- Demonstrable experience mentoring engineers, leading technical initiatives, or acting as a technical lead within a team.
- Strong communication skills, with the ability to document standards clearly and influence adoption across engineering teams.
What’s included:
As Lead DevOps Engineer at our award-winning headquarters in Wrexham, you’ll enjoy welcoming, spacious, state-of-the-art offices, plus communal spaces including a treehouse – we like to do things differently around here! You’ll also benefit from:
- Permanent contract
- 25 days annual leave plus bank holidays (pro rata)
- Mental health support (through our Employee Assistant Programme) with access to an on-site mental health counsellor
- Access to our wellbeing room to help enhance your physical and mental wellbeing
- Access to a 24/7 doctor line
- Subsidised meals
- Free on-site gym access
- And did we mention our epic parties? We know how to celebrate in style!
We want to hire the whole version of you:
If you require any adjustments to the recruitment process, please let us know so we can help you to be at your best!
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search