Infrastructure Engineer
Indexed description
Our Core Infrastructure Stack
OpenShift Platform & Virtualization
- OpenShift Container Platform (OCP) administration on bare-metal/on-prem environments
- OpenShift Virtualization for managing VM-based workloads alongside containers
- RHEL (Red Hat Enterprise Linux) advanced system administration and performance tuning
- OpenShift Networking (Software Defined Networking / OVN-Kubernetes)
- OpenShift Data Foundation (ODF) for persistent, software-defined storage
- Load balancing, ingress controllers, and DNS management within a restricted network
- Cluster Security (Security Context Constraints (SCC), RBAC, Network Policies)
- Monitoring & Alerting (Prometheus, Grafana, and Alertmanager)
- Disaster Recovery (DR) and high-availability (HA) planning for stateful workloads
- Knowledge of Infrastructure as Code (Ansible/OpenTofu) for cluster day-2 operations.
- Convert high-level blueprints into detailed cluster specs, including IP address management (IPAM), storage topology, and VLAN tagging for bare-metal OCP.
- Bare-Metal Administration: Independently lead the installation, scaling, and lifecycle management of OpenShift on physical hardware, including firmware updates and node provisioning.
- Storage Orchestration: Architect and manage OpenShift Data Foundation (ODF), ensuring persistent storage is highly available and performant for simulation databases.
- Network Engineering: Design and troubleshoot complex OVN-Kubernetes network patterns, ensuring secure and low-latency communication between simulation services.
- OpenShift Virtualization: Lead the migration and management of legacy VM workloads into OpenShift, providing a unified control plane for the digital twin ecosystem.
- Cluster Security Hardening: Implement and enforce strict SCC and RBAC policies, ensuring the platform complies with global security benchmarks and the squad's DevSecOps standards.
- Monitoring & DR: Establish comprehensive monitoring with Prometheus/Grafana and independently design and test Disaster Recovery scenarios to ensure zero-data-loss targets.
- Autonomous Troubleshooting: Act as the final escalation point for cluster-level performance issues, networking bottlenecks, or storage failures, resolving them without external guidance.
- 6 to 8 years of experience in Infrastructure Engineering or Systems Administration.
- Expert-level mastery of Red Hat OpenShift (OCP); Red Hat Certified Specialist in OpenShift Administration/Automation is highly preferred.
- Deep experience in Bare-Metal deployments and RHEL administration.
- Hands-on proficiency with OpenShift Data Foundation (ODF) and OpenShift Virtualization.
- Strong understanding of Kubernetes Networking (SDN/OVN) and security constructs (Network Policies).
- Demonstrated ability to work independently on complex hardware-to-software integration projects.
- Experience supporting high-compute Digital Twin or Simulation environments.
- Experience with backup solutions for OpenShift (Velero or OADP).
- Experience in managing on-premises infrastructure for global product engineering hubs.
- Exposure to industrial domains such as Manufacturing, Logistics, or Transportation is a plus.
- Exposure to CI/CD and platform automation tools such as GitHub Actions, Jenkins, or GitOps practices to better align with DevSecOps operating models.
- Hybrid-cloud exposure (AWS, Azure, or GCP) as a preferred qualification
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search