Linux & Kubernetes Administrator
Indexed description
Roles & Responsibilities
- Linux Infrastructure Administration
Manage system lifecycle activities, including provisioning, patching, upgrades, hardening, and decommissioning.
Monitor system health, availability, resource utilization, and performance metrics.
Manage user administration, access controls, and privilege management.
- Container Platform Administration
Manage Kubernetes clusters across development, testing, and production environments.
Configure and maintain Namespaces, Pods, StatefulSets, Deployments, Services, Ingress Controllers, and Persistent Volumes.
Ensure container platform availability, scalability, and performance.
- Cloud-Native Platform Management
Implement cloud-native operational best practices for containerized applications.
Manage container registries, image repositories, and artifact management solutions.
Support microservices-based application deployments.
Support AI/ML, GPU, and high-performance computing infrastructure.
Support GPU/ AI environment
- Platform Reliability & Performance Management
Conduct root cause analysis (RCA) for production incidents.
Perform proactive health checks, capacity planning, and resource optimization.
Implement monitoring and alerting for infrastructure and container workloads.
- Security & Compliance
Manage vulnerability remediation and security patching.
Configure security policies for Kubernetes and OpenShift environments.
Implement container image scanning, compliance monitoring, and runtime protection.
Ensure compliance with ISO 27001, PCI-DSS, SOC2, and organizational governance requirements.
- Automation & DevOps Enablement
Implement Infrastructure as Code (IaC) using Terraform, Ansible, and GitOps methodologies.
Support CI/CD pipelines integrated with Kubernetes and container platforms.
Drive operational automation and service reliability improvements.
- Backup, Recovery & Business Continuity
Support disaster recovery planning and testing.
Ensure platform resilience and adherence to RPO/RTO commitments.
Participate in business continuity and disaster recovery exercises.
- Monitoring & Observability
Configure dashboards, alerts, and operational reporting.
Analyze trends and recommend performance improvements.
Support observability initiatives across cloud-native platforms.
- Stakeholder & Service Management
Collaborate with Application, DevOps, Cloud Engineering, Security, and Infrastructure teams.
Support audits, compliance reviews, and customer governance meetings.
Mentor junior administrators and provide technical leadership.
Experience & Educational Requirement
BE/B-Tech or equivalent with Computer Science or Electronics & Communication
Relevant Experience
- Minimum 8-12 years of experience in Linux System Administration.
- Minimum 5+ years of hands-on experience supporting Kubernetes/OpenShift platforms in production environments.
- Experience managing large-scale enterprise and cloud-native environments.
- Strong exposure to container orchestration technologies.
- Experience supporting mission-critical 24x7 managed services operations.
- Experience in DevOps, CI/CD, Infrastructure as Code, and automation frameworks.
- Understanding of SRE principles and platform reliability engineering practices.
- Scripting knowledge in Python, Bash, or PowerShell.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search