Systems Operations and Engineering Manager
Indexed description
About the Opportunity
Our client is seeking an experienced Manager, Systems Engineering & Operations to lead the team responsible for the reliability, security, and operational excellence of the organization's enterprise systems and platforms. This leadership role oversees engineers supporting both on-premises and cloud infrastructure while ensuring production environments remain secure, resilient, highly available, and audit-ready.
This is an excellent opportunity for someone who enjoys building high-performing engineering teams, driving operational excellence, modernizing infrastructure, and improving system reliability through automation and continuous improvement. The ideal candidate is a collaborative leader who thrives in fast-paced environments and enjoys balancing day-to-day operational responsibilities with long-term platform strategy.
What You'll Do
- Provide technical leadership during complex production incidents and high-risk infrastructure changes.
- Lead incident response efforts, root cause analysis, and problem management while driving permanent solutions to recurring issues.
- Improve operational visibility through monitoring, dashboards, metrics, logging, and alerting.
- Lead, mentor, and develop a team of Systems Engineers responsible for supporting enterprise production platforms and services.
- Partner with Information Security to ensure systems remain secure, compliant, and audit-ready.
- Oversee patch management, access controls, change management, and regulatory compliance activities.
- Contribute to platform architecture decisions with a focus on operational scalability, reliability, and long-term sustainability.
- Support disaster recovery planning, business continuity testing, and technical debt reduction initiatives.
- Establish engineering standards and best practices for system lifecycle management, operational readiness, and platform reliability.
- Foster a culture of ownership and accountability for production systems from deployment through retirement.
- Oversee the operational health of enterprise systems, including compute, storage, identity management, and platform services.
- Ensure systems meet availability, performance, resiliency, and disaster recovery objectives.
- Champion automation initiatives, including Infrastructure as Code and automated operational processes, to improve efficiency and reduce manual effort.
- Manage vendor relationships, service providers, licensing, capacity planning, and infrastructure costs.
What We're Looking For
- Bachelor's degree in Information Technology, Systems Engineering, Computer Science, or a related field (or equivalent professional experience).
- 7+ years of experience supporting enterprise production systems and infrastructure.
- 2+ years of experience leading technical teams or serving as a senior technical lead.
- Strong background in systems operations, infrastructure engineering, site reliability engineering (SRE), or platform engineering.
- Experience supporting hybrid infrastructure environments, including both on-premises and cloud platforms.
- Strong knowledge of production operations, incident management, system reliability, and operational best practices.
- Experience working in security-focused or audit-sensitive environments.
- Familiarity with observability platforms, monitoring tools, logging, metrics, and alerting solutions.
- Experience implementing automation, Infrastructure as Code, configuration management, or operational tooling.
- Strong understanding of disaster recovery, business continuity, systems security, and operational risk management.
- Excellent communication, leadership, and collaboration skills with the ability to lead through complex technical challenges.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search