Back to search
Confidential DFW Linkedin · Posted 6d ago

Software Engineer - Fleet Automation

Dallas-Fort Worth Metroplex

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

RESPONSIBILITIES

  • Design, build, and maintain fleet automation services and internal platforms for provisioning, configuration, and lifecycle management of large-scale GPU and CPU compute nodes.
  • Develop APIs and service integrations that enable Infrastructure and Operations teams to deploy, image, validate, and decommission hardware with minimal manual intervention.
  • Build and maintain backend services in Go, C#, and TypeScript with a strong focus on reliability, testability, and long-term maintainability.
  • Design and evolve data models and persistent state for automation workflows, working across relational and NoSQL databases as appropriate.
  • Build and maintain CI/CD pipelines that gate configuration changes, run automated hardware validation tests, and promote changes safely across environments.
  • Instrument systems for observability — designing metrics, alerts, and dashboards in Prometheus and Grafana that provide real-time fleet health visibility to on-call teams.
  • Participate in on-call rotations; own incident response, post-mortems, and follow-through on reliability improvements across the fleet.
  • Identify systemic gaps in fleet reliability and efficiency and champion engineering solutions that reduce operational toil at scale.


REQUIREMENTS

  • Bachelor’s Degree in Computer Science, Software Engineering, or equivalent practical experience.
  • 5+ years of software engineering experience building production backend services or infrastructure automation tooling.
  • Proficiency in Go, C#, or TypeScript.
  • Experience designing and working with relational and NoSQL databases to support stateful automation workflows and internal platform services.
  • Solid understanding of Linux systems — networking, storage, process management, and debugging on Ubuntu or RHEL variants.
  • Experience building and maintaining CI/CD pipelines and observability stacks (Prometheus, Grafana, Alertmanager, ELK) in a production environment.
  • Familiarity with GPU compute infrastructure and NVIDIA tooling (DCGM, nvidia-smi, NVIDIA Container Toolkit) is a strong plus.
  • Exposure to event-driven architectures or messaging platforms (e.g. Kafka) is a plus for teams building automation workflows across distributed services.
  • Strong communication skills and a collaborative mindset — comfortable navigating ambiguity, taking initiative, and working across Infrastructure, Operations, and Research teams.


Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search