Service Excellence Engineer
Indexed description
Key responsibilities:
- Lead technical response during critical service incidents, ensuring swift recovery and minimal business disruption.
- Build early-warning and real-time visibility using observability platforms and monitoring data.
- Develop dashboards, alert thresholds, and recovery indicators for critical services and infrastructure.
- Conduct structured root-cause analysis and drive permanent corrective actions to prevent recurrence.
- Collaborate with Network, Platform, and Application teams to strengthen continuity measures and response readiness.
- Implement automation-driven remediation steps, reducing manual resolution time and repetitive interventions.
- Maintain and prioritise a continuity backlog focused on recurrence prevention and operational gaps.
- Reduce false alarms, repeated disruptions, and reactive firefighting through proactive engineering practices.
- Improve uptime, recovery time, fault prevention, and continuity metrics across SbM-supported services.
- Create runbooks, playbooks, and standard response workflows to enable consistent execution.
- Provide continuity insights and incident learnings to leadership, enabling preventive decisions.
- Champion proactive continuity thinking and operational preparedness across teams and regions.
- 5–10 years in platform support, infrastructure operations, critical incident response, or technical operations.
- Experience with monitoring systems, dashboards, event analysis, and service telemetry.
- Demonstrated incident handling and structured recovery experience in enterprise environments.
- Hands-on exposure to networks, cloud platforms, infrastructure components, or enterprise applications.
- Ability to automate repetitive tasks using scripts or operational workflows.
- Experience across distributed or multi-region operational environments preferred.
- Bachelor’s degree in computer science, Engineering, IT, or related technical field.
- Certifications in cloud, network, or operations (Azure/AWS/Network/IT Ops) are beneficial.
- Strong analytical and diagnostic capabilities with structured problem-solving skills.
- Effective communication during high-pressure incidents and cross-functional coordination ability.
- Familiarity with automation concepts, observability, and operational engineering practices is an advantage.
- Excellent communication skills with technical and non-technical stakeholders
- Keen sense of ownership and motivation
- Strong analytical skills
- Ability to work collaboratively with others
- Ability to multi-task
- Very strong skills in representing operational instructions and information in a clear and understandable way.
- An agile, proactive person who takes an attitude of ‘getting things done’ and delivers on objectives
- Ability to influence others at all levels.
- An ability to lead, syndicate and achieve organisational change
- Strong interpersonal skills
- Detail-oriented
- Ability to take initiative and work independently
- Results-oriented
- Sense of urgency
- Focused to deliver best in class service
- Energetic and enthusiastic combined with resilience and determination
- Ability to take initiative and work independently
- Results-oriented and a sense of urgency
- Bigger picture view
We are happy to support your need for any adjustments during the application and hiring process. If you need special assistance or an accommodation to use our website, apply for a position, or to perform a job, please contact us by emailing [email protected].
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search