Quality Engineering Manager
Indexed description
- Serve as a technical escalation point for critical production incidents, outages, and service degradation events.
- Lead troubleshooting and root cause analysis efforts across applications, integrations, infrastructure, cloud services, APIs, and supporting technologies.
- Coordinate incident response activities involving application teams, infrastructure teams, vendors, and business stakeholders.
- Restore service quickly while ensuring long-term corrective actions are identified and implemented.
- Participate in major incident management processes and post-incident reviews.
- Apply Site Reliability Engineering principles to improve platform reliability, scalability, resilience, and operational efficiency.
- Define and support Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational performance metrics.
- Drive reduction of operational toil through automation and process improvement.
- Support production readiness reviews and operational acceptance processes.
- Participate in disaster recovery, resiliency, failover, and business continuity testing.
- Analyze complex system behavior using logs, metrics, traces, performance data, and monitoring tools.
- Perform deep technical investigations across application, infrastructure, data, network, and cloud environments.
- Identify recurring issues, trends, and systemic problems to reduce future incidents.
- Lead root cause analysis (RCA) activities and implement preventive solutions.
- Develop technical recommendations that improve system stability, performance, and reliability.
- Design, implement, and optimize monitoring, alerting, logging, and observability solutions.
- Develop dashboards and health indicators providing visibility into application and platform performance.
- Partner with engineering teams to improve observability through instrumentation, distributed tracing, synthetic monitoring, and telemetry collection.
- Continuously refine alerting strategies to reduce false positives and alert fatigue.
- Establish operational health metrics and reliability reporting.
- Develop and maintain automation solutions that improve operational efficiency and service reliability.
- Create scripts, tools, and workflows to automate diagnostics, health checks, remediation activities, and routine support tasks.
- Leverage Infrastructure as Code (IaC) and automation frameworks where appropriate.
- Drive continuous improvement through operational automation and self-healing capabilities.
- Partner with engineering teams to integrate automation into deployment and operational workflows.
- PowerShell
- Python
- Bash/Shell
- SQL
- REST APIs
- Workflow automation platforms
- Leverage AI and Generative AI tools to improve incident analysis, troubleshooting, knowledge management, and operational efficiency.
- Utilize AI-powered operational insights to identify patterns, anomalies, and emerging risks.
- Contribute to development of intelligent support capabilities including chatbots, operational copilots, automated RCA generation, and knowledge recommendations.
- Evaluate opportunities to improve production support through AI-enabled automation and predictive analytics.
- Promote responsible AI practices aligned with enterprise governance and security requirements.
- Develop, maintain, and continuously improve support runbooks, operational procedures, troubleshooting guides, and recovery playbooks.
- Ensure support documentation remains accurate, actionable, and aligned with production environments.
- Establish standardized operational processes supporting incident response and service recovery.
- Capture lessons learned from incidents and incorporate improvements into support practices.
- Build and maintain operational knowledge repositories to improve support consistency and reduce resolution times.
- Support cloud-hosted and hybrid application environments, including Azure-based platforms and services.
- Assist engineering teams in implementing resilient and observable cloud architectures.
- Monitor cloud resource health, performance, utilization, and operational readiness.
- Support cloud deployments, platform upgrades, and operational change activities.
- Partner with cloud engineering teams on modernization and reliability initiatives.
- Partner closely with Engineering, Product, Architecture, Infrastructure, Security, QA, and Vendor teams.
- Review operational readiness of new systems and platform enhancements.
- Mentor junior support engineers and provide technical guidance during incident response activities.
- Promote operational excellence, reliability engineering principles, and continuous improvement practices.
- Serve as a subject matter expert within assigned technology domains.
- Ensure production support activities comply with enterprise risk, security, regulatory, and audit requirements.
- Identify operational risks and escalate issues appropriately.
- Support implementation of internal controls and operational governance standards.
- Participate in audit, compliance, and regulatory review activities as required.
- Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent experience.
- Minimum 5+ years of experience supporting enterprise applications, platforms, or infrastructure in production environments.
- Experience troubleshooting complex application, infrastructure, integration, or cloud-related issues.
- Strong knowledge of incident management, production support processes, and operational best practices.
- Experience with monitoring, observability, logging, and alerting platforms.
- Experience with scripting and automation technologies.
- Strong analytical, troubleshooting, and problem-solving skills.
- Excellent communication and collaboration skills.
- Experience with Site Reliability Engineering (SRE) practices.
- Experience defining and managing SLIs, SLOs, and operational metrics.
- Experience supporting distributed systems, APIs, microservices, and cloud-native applications.
- Experience performing production readiness reviews and operational assessments.
- Dynatrace
- Splunk
- Datadog
- Azure Monitor
- Grafana
- Prometheus
- OpenTelemetry
- AppDynamics
- Experience supporting Azure cloud environments.
- Familiarity with Azure App Services, AKS, Functions, Storage, Event Hub, Service Bus, and monitoring services.
- Understanding of cloud security and operational best practices.
- Advanced PowerShell scripting.
- Python development and automation.
- REST API integration.
- Infrastructure as Code concepts.
- CI/CD tools and deployment pipelines.
- Experience using Microsoft Copilot, Azure AI services, Microsoft Foundry, or similar AI-enabled platforms.
- Familiarity with AI-assisted troubleshooting and operational analytics.
- Experience implementing AI-enabled support workflows or knowledge management solutions.
- Understanding of intelligent automation and operational copilots.
- Resolves complex incidents quickly and effectively.
- Proactively identifies and eliminates recurring production issues.
- Uses automation and scripting to eliminate manual support work.
- Builds comprehensive support playbooks and operational runbooks.
- Leverages AI to accelerate troubleshooting and operational insights.
- Establishes strong observability and monitoring capabilities.
- Partners effectively with engineering teams to improve reliability.
- Drives a culture of operational excellence through SRE principles.
- Continuously improves customer experience through system stability, resilience, and performance.
Location
Buffalo, New York, United States of America
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search