Mercedes AMG High Performance Powertrains
Linkedin · Posted 14d ago
HPC Engineer
Continue to application
Add your email once, then Caio opens the original posting.
Indexed description
Purpose of the role is to…
- Own the reliability, availability and performance of HPC platforms supporting simulation, analysis and engineering workloads
- Improve compute, storage, networking and scheduling services to enable efficient, scalable workload delivery
- Provide technical escalation for HPC incidents, capacity issues, performance bottlenecks and complex user problems
- Administering Linux-based HPC clusters, including compute nodes, schedulers and shared platform services
- Troubleshooting issues across hardware, OS, network, storage, applications and user workflows
- Managing capacity, performance and availability for engineering and simulation workloads
- Automating operational tasks using Bash, Python, PowerShell, Ansible or equivalent tools
- Translating technical user requirements into practical service improvements
- Supporting Linux-based HPC, scientific computing, simulation or high-throughput compute environments
- Diagnosing workload, queue, licence, performance, data movement and application issues
- Operating at a senior technical level in an enterprise or engineering-led environment
- Delivering maintenance, upgrades, patching and change activity with minimal service impact
- Working with suppliers and internal teams to resolve platform issues and improve service maturity
- HPC architecture, parallel workloads, scheduling, queues and resource allocation
- Linux administration, scripting, patching and secure configuration
- Schedulers such as Slurm, PBS, LSF or equivalent
- Scale-out storage, file systems, backup, archive and data lifecycle management
- Networking, interconnects, latency, bandwidth and data locality considerations
- Monitoring, performance tuning, benchmarking and capacity forecasting
- Security, vulnerability management, access control and compliance for shared platforms
- Desirable: motorsport, automotive, CFD, simulation or data science experience
- Relevant degree, apprenticeship, professional qualification or equivalent technical experience
- Relevant technical certifications, or equivalent experience, in Linux, HPC, storage, networking, automation or ITIL
- Analytical, curious and comfortable solving complex technical problems
- Proactive in improving resilience, reducing risk and removing operational friction
- Structured, communicative and effective across hands-on delivery and change control
- Collaborative, customer-focused and willing to share knowledge
- Maintain stable, secure and performant HPC services for critical engineering workloads
- Improve compute and storage utilisation through effective monitoring, queue management and capacity planning
- Resolve incidents quickly and reduce repeat issues through automation, documentation and service improvement
- Deliver upgrades, maintenance and project work safely with clear communication and change control
- Improve simulation throughput, data availability and user productivity
- Define and guide strategic direction on HPC related topics
- The role combines operational support and project delivery, including planned maintenance, capacity improvement, lifecycle management and occasional out-of-hours activity
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search
Want help applying to roles like this?
Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search