Senior OpenShift / Kubernetes Engineer
Indexed description
Responsibilities
Key Responsibilities
- Administer and support Red Hat OpenShift and Kubernetes environments.
- Manage and troubleshoot application deployments across OpenShift/Kubernetes clusters.
- Configure and maintain Jenkins deployments and CI/CD pipelines.
- Configure Jenkins pipeline jobs, deployment parameters, and automation workflows.
- Perform production troubleshooting across pods, deployments, services, routes, storage, networking, and cluster resources.
- Handle P1/critical production incidents, perform root-cause analysis, and coordinate resolution within defined SLAs.
- Perform OpenShift/Kubernetes cluster upgrades and validate cluster health after upgrades.
- Manage and perform Operator upgrades, including validation and troubleshooting of Operator-related issues.
- Monitor and maintain overall cluster health, including nodes, workloads, resource utilization, operators, storage, and networking.
- Troubleshoot issues during deployment migrations between environments/clusters.
- Understand the complete application and infrastructure environment to identify dependencies and troubleshoot issues effectively.
- Analyze logs, events, pod status, resource utilization, and deployment configurations to identify root causes.
- Collaborate with application, middleware, DevOps, and infrastructure teams during deployments and production incidents.
- Follow change-management, incident-management, and operational procedures.
- Prepare technical documentation, troubleshooting guides, and post-incident RCA documents.
- 6+ years of experience in DevOps / Kubernetes / OpenShift administration.
- Strong hands-on experience with Red Hat OpenShift.
- Good knowledge of Kubernetes concepts and administration.
- Hands-on experience with Jenkins deployments and pipeline configuration.
- Strong production troubleshooting and incident-management skills.
- Experience handling P1 / critical production incidents.
- Experience with OpenShift/Kubernetes cluster upgrades.
- Experience with Operator installation, management, and upgrades.
- Good understanding of pods, deployments, StatefulSets, Services, Routes/Ingress, ConfigMaps, Secrets, PVCs/PVs, RBAC, nodes, and namespaces.
- Good understanding of application deployment and migration processes.
- Strong understanding of cluster health monitoring and capacity/resource management.
- Good knowledge of Linux and command-line troubleshooting.
- Experience with enterprise production OpenShift environments.
- Knowledge of CI/CD and DevOps practices.
- Experience with Kafka or other middleware platforms is an advantage.
- Knowledge of monitoring and logging platforms such as Prometheus, Grafana, ELK, or equivalent.
- Strong analytical, troubleshooting, communication, and incident-management skills.
OpenShift/Kubernetes Administration | Jenkins & CI/CD | Production Troubleshooting | P1 Incident Management | Cluster Upgrades | Operator Upgrades | Deployment Migration | Cluster Health Management | Root Cause Analysis
Requirements
Key Responsibilities
- Administer and support Red Hat OpenShift and Kubernetes environments.
- Manage and troubleshoot application deployments across OpenShift/Kubernetes clusters.
- Configure and maintain Jenkins deployments and CI/CD pipelines.
- Configure Jenkins pipeline jobs, deployment parameters, and automation workflows.
- Perform production troubleshooting across pods, deployments, services, routes, storage, networking, and cluster resources.
- Handle P1/critical production incidents, perform root-cause analysis, and coordinate resolution within defined SLAs.
- Perform OpenShift/Kubernetes cluster upgrades and validate cluster health after upgrades.
- Manage and perform Operator upgrades, including validation and troubleshooting of Operator-related issues.
- Monitor and maintain overall cluster health, including nodes, workloads, resource utilization, operators, storage, and networking.
- Troubleshoot issues during deployment migrations between environments/clusters.
- Understand the complete application and infrastructure environment to identify dependencies and troubleshoot issues effectively.
- Analyze logs, events, pod status, resource utilization, and deployment configurations to identify root causes.
- Collaborate with application, middleware, DevOps, and infrastructure teams during deployments and production incidents.
- Follow change-management, incident-management, and operational procedures.
- Prepare technical documentation, troubleshooting guides, and post-incident RCA documents.
- Experience with enterprise production OpenShift environments.
- Knowledge of CI/CD and DevOps practices.
- Experience with Kafka or other middleware platforms is an advantage.
- Knowledge of monitoring and logging platforms such as Prometheus, Grafana, ELK, or equivalent.
- Strong analytical, troubleshooting, communication, and incident-management skills.
OpenShift/Kubernetes Administration | Jenkins & CI/CD | Production Troubleshooting | P1 Incident Management | Cluster Upgrades | Operator Upgrades | Deployment Migration | Cluster Health Management | Root Cause Analysis
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search