OpenShift SRE
Indexed description
This role is responsible for OpenShift platform deployment, BareMetal Day 0 / Day 1 setup support, platform troubleshooting, GitOps-based deployment support, production incident stabilization, escalation support, handoff documentation, and operational continuity. The consultant must be able to operate independently after onboarding, troubleshoot issues under pressure, and support service restoration during overnight incidents.
The ideal candidate has hands-on experience deploying OpenShift on BareMetal servers, supporting OpenShift GitOps / ArgoCD workflows, using Ansible to automate OpenShift or Linux operations, and troubleshooting production issues across clusters, applications, deployments, nodes, operators, routes/ingress, networking, storage, and platform services.
Core Responsibilities:
- Support BareMetal OpenShift deployment and Day 0 / Day 1 platform setup activities where applicable.
- Support OpenShift deployment prerequisites and troubleshooting across DNS, load balancers, bootstrap/control plane/worker node setup, RHCOS, ignition, API/API-int/ingress, firewall rules, and registry/mirror dependencies.
- Support OpenShift GitOps / ArgoCD deployment patterns, including failed sync troubleshooting, deployment validation, rollback coordination, and drift prevention.
- Use Ansible and automation-driven processes to support repeatable OpenShift operational tasks.
- Provide Tier 3 OpenShift SME support during assigned coverage windows.
- Support escalated production issues within the T-Mobile OpsOne OpenShift environment.
- Troubleshoot OpenShift platform and application issues across clusters, workloads, nodes, operators, routes/ingress, networking, storage integrations, and platform services.
- Support OpenShift deployment activities, including GitOps-based deployment patterns where applicable.
- Diagnose issues from limited initial information such as alerts, observability signals, failed GitOps syncs, BareMetal deployment issues, one-line incident reports, or escalations from another team.
- Help stabilize impacted services and restore platform or application availability during production incidents.
- Determine whether issues are application-related, platform-related, BareMetal infrastructure-related, networking-related, storage-related, operator-related, GitOps/deployment-related, or require escalation.
- Provide recommended resolution paths and communicate clearly during active incidents.
- Document actions taken, current status, open risks, and handoff notes.
- Provide RCA inputs where applicable and reasonably available.
- Support runbook updates, operational documentation, and backlog remediation during lower-volume periods.
- Participate in rotational support model with strong cross-shift handoffs.
- Rotate through day-shift alignment with Red Hat team to stay current on program context, deployments, known issues, and operational changes.
- Build trust quickly with Red Hat and T-Mobile stakeholders through calm, clear, and reliable communication.
Technologies
- Hands-on experience deploying OpenShift on BareMetal servers, including full platform setup and Day 0 / Day 1 operations.
- Experience with BareMetal OpenShift deployment prerequisites such as DNS, load balancers, bootstrap node, control plane and worker nodes, RHCOS, ignition configs, install-config.yaml, API / API-int / ingress endpoints, pull secrets, SSH keys, NTP, firewall rules, and registry or mirror configuration where applicable.
- Strong experience with OpenShift GitOps, preferably ArgoCD, including sync health, failed sync troubleshooting, deployment validation, rollback, and avoiding configuration drift.
- Strong hands-on experience with GitOps and CI/CD tooling such as OpenShift GitOps, ArgoCD, Helm, Kustomize, Tekton, Jenkins, GitLab, or similar tools. Candidate must be able to explain failed deployment troubleshooting and recovery without creating drift between Git and the live cluster.
- Experience using Ansible to automate OpenShift or Linux operations, including repeatable operational tasks, configuration changes, platform validation, health checks, user/RBAC setup, patching, or deployment support.
- Experience with Linux/RHEL administration and troubleshooting in support of OpenShift environments.
- Familiarity with observability, monitoring, logging, and alerting tools such as Prometheus, Grafana, Splunk, Dynatrace, or similar platforms.
- Experience working within enterprise ticketing, escalation, and collaboration tools such as ServiceNow, Jira, Slack, Microsoft Teams, or similar systems.
- Knowledge of Kubernetes concepts, including pods, deployments, services, config maps, secrets, namespaces, RBAC, persistent volumes, and networking.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search