Senior Platform Engineer – VMware & Kubernetes
Indexed description
Overview / Summary
We are seeking a Senior Platform Engineer with strong experience in virtualization, Kubernetes, cloud infrastructure, automation, scripting, and platform support. This role is responsible for engineering, deploying, administering, securing, and supporting enterprise virtualization platforms and related infrastructure.
The position will focus on VMware and OpenShift Virtualization (OSV), capacity management, automation, observability, troubleshooting, solution design, technical consulting, documentation, and platform support.
Key Responsibilities
- Engineer, deploy, administer, and protect enterprise virtualization solutions supporting infrastructure systems, business applications, and third-party applications.
- Architect, design, install, and administer VMware and OpenShift Virtualization (OSV) environments.
- Manage the full lifecycle of infrastructure technologies, including security vulnerability patching, planning, design, implementation, maintenance, upgrades, and decommissioning.
- Engineer, test, and document procedures, monitoring, logging, disaster recovery processes, security policies, and guidelines.
- Support large-scale deployments of virtualization technologies.
- Support HPE Synergy/ProLiant iLO and firmware field testing.
- Conduct capacity planning and forecasting for compute/virtual machines, memory, storage, and network resources.
- Analyze resource utilization trends and recommend infrastructure scaling, consolidation, or optimization.
- Collaborate with application teams and stakeholders to understand future demand and project capacity requirements.
- Develop and maintain capacity models and reports to support strategic planning.
- Develop automation solutions, including scripts and playbooks, for repetitive VMware/OSV tasks such as configuration changes, VM management, auditing, remediation, and ticketing system integration.
- Implement automation to efficiently deliver operator updates and changes at scale.
- Apply Site Reliability Engineering (SRE) principles and practices to improve platform stability, performance, and operational efficiency.
- Deploy and audit role-based access controls.
- Manage namespaces and resource quotas for CPU, disk, and storage.
- Implement and maintain end-to-end observability solutions, including monitoring, logging, and tracing for VMware/OSV environments.
- Integrate observability solutions with tools such as Dynatrace, RHACM, and Prometheus/Grafana.
- Explore and implement Event Driven Architecture (EDA) for enhanced real-time monitoring and response.
- Develop capabilities to identify and report abnormalities and observability gaps.
- Perform root cause analysis and troubleshoot issues across the global compute environment.
- Monitor virtual machine health, resource usage, and performance metrics proactively.
- Monitor for unusual activity that may indicate a compromise or misconfiguration.
- Provide technical consulting and expertise to application teams requiring VMware/OSV solutions.
- Design, implement, and validate custom or dedicated OSV clusters and VM solutions for applications with unique or complex requirements.
- Create, maintain, and update internal documentation and customer-facing content to support self-service and clearly communicate platform capabilities.
- Provide L1–L3 support to Operations teams for environment-related issues.
- Participate in monthly after-hours and weekend support as required.
- 10+ years of IT experience.
- 8+ years of development experience.
- Practical experience in two coding languages or advanced practical experience in one coding language.
- Understanding of VMware and Kubernetes concepts.
- Experience with Linux administration and networking fundamentals.
- Proficiency in scripting languages for automation.
- Experience with monitoring tools and logging solutions.
- Understanding of virtualization concepts and technologies, including KVM and VMware.
- Strong problem-solving and troubleshooting skills across multiple layers of the technology stack.
- Knowledge of CI/CD pipelines and DevOps methodologies.
- Experience with scripting and automation.
- Experience with cloud architecture and cloud infrastructure.
- Experience with VMware and VMware ESX Servers.
- Experience with platform support and infrastructure architectures.
- Experience with GitHub.
- Knowledge of change management, technical analysis, and IT solutions.
- Strong communication and collaboration skills.
- Self-starter with the ability to identify opportunities to evolve services.
- Associate degree or college senior status.
- Ansible
- GCP
- Dynatrace
- PowerShell
- Access controls
- Python
- Information security
- Automation
- Artificial intelligence and expert systems
At HTC Global Services, our employees have access to a comprehensive benefits package. Benefits can include Group Health (Medical, Dental, and Vision), Paid Time Off, Paid Holidays, 401(k) matching, Group Life and Disability insurance, Professional Development opportunities, Wellness programs, and a variety of other perks.
Our success as a company is built on inclusion and diversity. HTC Global Services is committed to providing a workplace free from discrimination and harassment, where every employee is treated with dignity and respect. We celebrate differences and believe that diverse cultures, perspectives, and skills drive innovation and success. HTC is an Equal Opportunity Employer and a proud National Minority Supplier. We seek to empower each individual, fostering an environment where everyone feels valued, included, and respected.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search