Back to search
Auto Hauler Exchange Linkedin · Posted today

Platform Engineer - Site Reliability

Rochester

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

ABOUT US:

Auto Hauler Exchange (AHX) is an innovative startup revolutionizing the auto transport and logistics industry. Our platform connects vehicle shippers and carriers to streamline vehicle transportation with real-time tracking, transparent pricing, and an easy-to-use interface. As we continue to grow, Auto Hauler Exchange is looking for an experienced Platform Engineer - Site Reliability to join our software development team in Rochester, MI.


JOB OVERVIEW:

AHX runs a federated GraphQL platform, a set of Go services, and several Next.js applications on Google Cloud. This is the first role here whose whole job is the ground they run on. You'll own our environments, our delivery pipeline and the quality gates in it, and our observability, and you'll build them as a product the rest of engineering can use without thinking about it. The surface area is wide and the autonomy is real. Quality is part of that ground, not a separate initiative. We want automated gates carrying the weight, so that a green pull request means something specific rather than meaning nobody looked: test suites and coverage, static analysis, license and dependency compliance, container scanning. You'll also shape where our pipelines go from here, including where GitHub Actions is the right tool and where a dedicated platform earns its place.


KEY RESPONSIBILITIES:

  • Own our Google Cloud environments: GKE, Cloud Run, and everything around them.
  • Define our infrastructure in Terraform, with drift detection so the code and the cloud stay in agreement.
  • Standardize CI/CD on reusable workflows, so a pipeline improvement lands everywhere at once.
  • Build quality gates into every pipeline, so a merge is blocked on objective signals rather than on someone remembering to look: unit and integration suites, coverage thresholds, static analysis, license and dependency compliance, and container scanning.
  • Own our code quality and security scanning toolchain across SAST, SCA, license compliance, and code quality. SonarQube, Checkmarx, FOSSA, Snyk, and their equivalents are all in scope, and part of the job is deciding which of them we actually need.
  • Define what green means for a pull request, and keep the gate fast enough that engineers do not route around it. A slow gate is a gate that gets skipped.
  • Inform our CI/CD strategy together with the engineering team: where GitHub Actions is the right tool, and where a dedicated platform such as Jenkins, Tekton, or Argo serves us better. Quality and automation are the deciding criteria, not familiarity.
  • Turn noisy scanner output into findings people act on, with severity thresholds, suppression policy, and clear ownership for what a gate blocks.
  • Build once and promote the same immutable artifact from development through staging to production, with configuration injected at deploy.
  • Automate release coordination across our shared internal SDK packages and the services that depend on them. • Get a correlation ID flowing through the GraphQL router, our subgraphs, our Go services, and our partner calls, so a single request can be traced start to finish.
  • Roll out OpenTelemetry tracing, error monitoring, and dashboards defined as code. • Set SLOs on the paths that make us money, and build alerting worth waking up for, with a runbook behind every alert.
  • Lead incidents and postmortems, and make sure the follow-up work actually ships.
  • Own secrets, workload identity, and least-privilege IAM.
  • Use AI tooling to move fast on the mechanical work: module authoring, workflow consolidation, runbook drafts, drift checks. Production changes still get a careful human review.
  • Track and optimize our cloud spend, and make the cost of a change visible before we ship it.
  • Mentor our junior engineer into owning monitoring, alerting, and runbooks.


SKILLS & QUALIFICATIONS:

  • A minimum of 3 years of experience in a DevOps, platform engineering, or site reliability role, with production ownership rather than project work.
  • Managed Kubernetes in production (GKE, EKS, or AKS).
  • Terraform in earnest. You've written modules, managed state, and cleaned up drift somebody else left behind.
  • GitHub Actions across more than one repository, including reusable workflows or composite actions.
  • Multi-stage Docker builds, and a promotion pipeline you've actually run.
  • Hands-on experience building automated quality gates into CI: test suite orchestration, coverage thresholds, static analysis, and security and license scanning.
  • Working experience with code quality and security scanning platforms such as SonarQube, Checkmarx, FOSSA, Snyk, or comparable SAST, SCA, and license compliance tooling.
  • Experience with more than one CI/CD platform. GitHub Actions plus Jenkins, Tekton, Argo, GitLab CI, CircleCI, or similar, and the judgment to say which one fits a given problem.
  • Observability as a practice: structured logging, distributed tracing, metrics, and SLOs with error budgets.
  • You've been on call for something that mattered, and written a postmortem that changed something.
  • Enough Go, Python, TypeScript, or Bash to read service code and know what to instrument.
  • Experience leveraging AI tools such as Claude, Cursor, or Copilot for infrastructure and automation work, with the review discipline to match.
  • Hands-on experience standing up BigQuery as an analytics platform: dataset and table design, partitioning and clustering, streaming and batch ingestion, and query cost controls.
  • Experience with the GCP data stack: Datastream or comparable replication, Pub/Sub subscriptions into BigQuery, Dataform or dbt for transformation, and a catalogue and lineage tool such as Dataplex.
  • Experience implementing warehouse governance: row-level or column-level security, IAM scoping, retention and deletion paths, and quotas set before a query surface opens up.
  • A bachelor's degree in Computer Science or a relevant field is preferred, but not required.


PREFERRED QUALIFICATIONS:

  • GCP specifically: Pub/Sub, Secret Manager, Artifact Registry, Workload Identity.
  • GraphQL federation in CI, or release engineering for versioned internal packages.
  • Experience defining SLOs and error budgets from scratch rather than inheriting them.
  • Experience migrating between CI/CD platforms, or consolidating several onto one.
  • Experience taking a noisy scanner and turning it into a gate teams trust rather than bypass.
  • Freight, logistics, or another domain where downtime has physical consequences.


How to Apply:

If you're ready to take on a diverse and exciting role in a fast-paced startup, we'd love to hear from you! Please submit your resume, along with a cover letter detailing your relevant experience and why you're interested in joining Auto Hauler Exchange.


Auto Hauler Exchange is an equal-opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.


Job Type: Full-time


Benefits:

  • 401(k)
  • 401(k) matching
  • Dental insurance
  • Employee assistance program
  • Health insurance
  • Health savings account
  • Life insurance
  • Paid time off
  • Retirement plan
  • Vision insurance


Schedule:

  • 8 hour shift
  • Weekends as needed


Work Location: Hybrid

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search