Back to search
Esconet Technologies Linkedin · Posted 14d ago

HPC Engineer

India

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description


Company: Esconet Technologies

Location: Delhi/NCR – Okhla Phase 1, New Delhi (On-site)

Experience: 5+ Years


About the Role


Esconet Technologies is looking for a highly motivated HPC System Integrator/Administrator with a strong passion for cluster administration, system integration, and validation of HPC clusters. In this role, you will influence the overall integration, delivery, and management of largely open-source HPC products and solutions — spanning Intel and AMD processors, NVIDIA GPUs, storage, InfiniBand, and Linux software.

A solid understanding of parallel processing (problem decomposition and work distribution), parallel programming (MPI, OpenMP), and computer architecture is essential for this role.


Minimum Qualifications


• Bachelor's degree in Computer Science, Computer Engineering, Computational Science, equivalent mathematical sciences, or a related field, with 3+ years of relevant work experience; or an equivalent combination of education, training, and experience

• 3+ years of experience with software development in Linux

• 3+ years of experience with HPC clusters and systems integration

• Ability to manage the AI stack up to the framework level

• Experience with GPU cluster capabilities and storage integration (LUSTRE/BeeGFS, cluster storage)

• Working knowledge of object-based storage, networking, InfiniBand, solution designing, and technical documentation


Key Responsibilities


• Install, configure, fine-tune, and troubleshoot multi-vendor, multi-site Linux HPC servers

• Build and deploy open-source software as well as vendor/partner software

• Diagnose and resolve system operational issues quickly and effectively

• Verify full operation of systems, including network, systems, and storage performance

• Configure scheduling and queuing systems

• Assist technical support teams with questions and issues encountered by customers

• Coordinate with vendors to resolve hardware and software problems

• Document system administration procedures for routine and complex tasks (wikis)

• Maintain and monitor the security of HPC systems and servers

• Design and configure HPC/Kubernetes clusters with NVIDIA GPU support, including MIG (Multi-Instance GPU) and AI Factory deployments

• Deploy and manage cluster file systems such as Lustre, including OSS (Object Storage Server) and MDS (Metadata Server) node configuration

• Build and maintain effective working relationships with coworkers, managers, and clients

• Travel as needed for on-site cluster installation or maintenance (limited)


Desired Skills & Experience


• Building, configuring, and administering Linux distributions — Rocky Linux, Ubuntu, CentOS, RHEL, and SUSE

• Windows Server administration

• Expert knowledge of parallel/distributed file systems such as Lustre or IBM GPFS

• Strong knowledge of networking and cluster-based distributed computing, including InfiniBand switch configuration

• Experience deploying open-source and commercial HPC platforms

• HPC cluster architecture design and configuration; AI cluster and Kubernetes cluster experience

• Strong scripting skills — Bash, Python, Perl; Ansible for automation is a plus

• Experience building dashboards (e.g., Flask-based) for cluster monitoring is a plus

• Skilled in diagnosing and debugging complex HPC hardware/software issues, with proposed workarounds

• Strong troubleshooting and root-cause analysis capabilities


Key Skill Areas at a Glance


  • Cluster Application & Design: HPC Cluster Application, HPC Cluster Architecture Designing, HPC Cluster Configuration
  • AI/GPU Infrastructure: AI Cluster, Kubernetes Cluster, NVIDIA GPU, MIG, AI Factory
  • Storage: Cluster File System (Lustre), Cluster Storage (OSS & MDS Nodes), BeeGFS, Object-Based Storage
  • Networking: InfiniBand, InfiniBand Switch, Networking & Solution Design
  • Automation/Scripting: Python, Bash, Perl, Ansible, Flask Dashboards
  • OS Administration: Rocky Linux, Ubuntu, CentOS, RHEL, SUSE, Windows Server


Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search