Back to search
Hoonify Technologies Inc. Linkedin · Posted 3d ago

AI Infrastructure Engineer - Inference Platform

Albuquerque

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

We are seeking an AI Infrastructure Engineer to help build, deploy, and operate the LLM serving infrastructure underpinning our AI/ML platform. This role focuses on implementation, automation, and optimization of production inference systems, tuning model serving runtimes for performance and scale across NVIDIA and AMD GPU fleets, working under the technical direction of senior engineering leadership and the platform's established architectural patterns.


You'll ship well-engineered, well-tested infrastructure changes and grow your depth in GPU-backed workloads, distributed model serving, observability, and continuous delivery. You'll work directly with senior engineers on real production systems, receive code and design review on everything you ship, and have a clear path to expanded scope and ownership as your experience deepens.


What You'll Do

  • Deploy, tune, and optimize high-performance LLM inference pipelines on GPU infrastructure, improving throughput, latency, and cost efficiency within established design patterns.
  • Analyze, profile, and optimize model serving workloads across inference frameworks such as vLLM, SGLang, and TensorRT-LLM, and across different model families and hardware architectures.
  • Build and operate scalable, production-grade API services for model inference, including request routing, multi-tenant isolation, usage metering, and observability.
  • Develop benchmarking harnesses, monitoring infrastructure, and automation tooling that make serving performance measurable and reproducible.
  • Scale inference workloads across multi-GPU, multi-node environments spanning NVIDIA and AMD accelerators.
  • Evaluate, prototype, and integrate model fine-tuning workflows and frameworks.
  • Collaborate closely with engineering and product teams to align infrastructure capabilities with customer-facing services.
  • Investigate and resolve issues across the stack, including container, node, network, and accelerator-level problems, escalating appropriately when scope exceeds the role.
  • Write clear documentation, including runbooks, internal references, and design notes for the changes you ship.
  • Participate in code and design reviews, both as author and reviewer, and incorporate feedback from senior engineers into your work.


Required Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, Applied Math, or Data
  • Science, plus three (3) years relevant work experience or equivalent combination of education and relevant experience.
  • Professional software engineering experience, with at least some of it touching ML systems, GPU workloads, or high-performance backend services.
  • Working knowledge of Kubernetes in a production context, including writing and
  • debugging manifests, understanding core resource types, and operating production workloads.
  • Hands-on experience serving or deploying LLMs — you've run vLLM, SGLang, TGI, TensorRT-LLM, or similar.
  • Comfort working in a Linux environment and with standard developer tooling, including Git-based workflows.
  • Familiarity with CI/CD systems and the basic mechanics of automated build, test, and deployment pipelines.
  • Strong proficiency in Python with familiarity at least one programming or scripting language used for infrastructure work (Go, Rust, C++ or Bash).


Preferred Qualifications

  • Experience building or operating retrieval-augmented generation (RAG) pipelines, including vector databases, embedding models, and retrieval serving at scale.
  • AMD/ROCm experience.
  • Experience cleaning and curating datasets for LLM training and fine tuning.
  • Fine-tuning experience of any depth: LoRA/QLoRA, full fine-tunes, dataset curation, or evaluation design.
  • Experience with usage metering, billing systems, or multi-tenant API platforms.
  • Experience instrumenting services and consuming observability data, including writing Prometheus queries, building Grafana dashboards, or working with distributed traces.
  • Experience with HPC batch schedulers and MPI based workloads.
  • Experience with alternative compute architectures for inference (RISC-V, FPGA, ASIC).


Why Join Hoonify

You'll have a direct line to leadership and genuine influence over the company's growth trajectory. This is a rare opportunity to build a cutting-edge multi-cloud computational platform at a company doing meaningful work in AI — with the autonomy and resources to make it your own.


About Our Team

Hoonify delivers secure, sovereign AI infrastructure designed for the next generation of inference workloads. Powered by TurbOS , our platform enables organizations and NeoCloud/data center operators to transform CPU/GPU infrastructure into production-ready AI environments—supporting local LLMs, agentic copilots, RAG, and embeddings. We empower teams with robust model lifecycle management, multi-tenant controls, usage metering, and fully auditable operations.


Hoonify is an equal opportunity employer. We welcome applicants from all backgrounds and are committed to building a diverse and inclusive team.


Must be eligible to obtain and maintain a US government security clearance.

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search