Platform / DevOps Engineer
Indexed description
WasteFlow runs computer-vision sensors on waste-processing lines — a fleet of on-prem, PLC-connected edge devices that stream detections and drive automation (conveyor control, alarms, sorting). Behind them sits a TypeScript product, a Postgres/TimescaleDB system of record, and a full observability and infrastructure-as-code platform on AWS.
We're hiring a Senior Platform / DevOps Engineer to own that platform end to end. This is a role with a single, clear owner: you will run our cloud infrastructure, our edge-fleet deployment system, our CI/CD, our databases, and our observability stack — and set the standards the rest of the team builds against. You'll keep production stable and make shipping dramatically easier — and you'll do it by building systems the whole team can run: documented, reproducible, and never dependent on any single person, yourself included.
Your main responsibilities are DevOps and fullstack. On top of that, you'll need a strong working understanding of MLOps — enough to own its plumbing (the MLflow model registry, artifact tracking, and how trained models get provisioned and pushed to sensors in the field) and to partner effectively with the ML team, who own the model logic itself.
What you'll ownCloud infrastructure & IaC. Our AWS footprint managed as code with OpenTofu; GitHub repositories and org scaffolding provisioned with Pulumi; sensor and host configuration automated with Ansible. You'll keep infrastructure reproducible, versioned, and reviewable — no snowflakes.
Product & fullstack. Our TypeScript product surface — a Next.js (React 18) frontend with Mantine, and Fastify services with zod contracts and Supabase/Postgres access. Alongside the platform, you'll build and maintain features across the stack, from API endpoints to UI, and keep the product codebase healthy. (Hena owns the fullstack side with you; you'll share this scope, not carry it alone.)
Edge fleet deployment. Our "MLHub / Control Center" pipeline — CLI and API driving an Ansible-orchestrated flow (Configure → Install → Provision → Deploy → Integrate → Initialize) out to devices in the field, pulling from GitHub, the model registry, and config stores. Edge software ships as versioned Python packages (built and published by CI, installed and updated on devices by Ansible) rather than containers. You'll harden this into reliable, observable, one-command provisioning and over-the-air updates across a growing fleet, connected over a Tailscale mesh.
CI/CD & developer platform. Our multi-repo setup (a core repo plus standardized sub-repos) and the automated workflows that guard it: format/lint, static analysis, automated tagging, test, and package publishing — versioned Python packages to our internal registry are the primary release artifact for edge software. You'll own repo standardization (pyproject.toml, packaging and versioning config, Dockerfiles for cloud services, editor and CI config), keep pipelines fast and green, and make "create a new service" a paved road.
Cloud compute & containers. Our cloud services run as Docker containers on AWS ECS / Fargate — build, ship, run, scale, and roll back predictably. (Edge devices are not containerized; they run versioned Python packages, as above.)
Databases & storage. PostgreSQL 17 with the TimescaleDB extension for sensor and detection time-series, on AWS RDS, plus S3 object storage. You'll own reliability, backups, migrations, access, and performance as data volume grows.
Observability & reliability. Our Grafana Cloud stack: Prometheus metrics, Loki logs, Grafana Alloy and Fluentd collection, Grafana Faro for frontend, CloudWatch, and K6 load testing — with service discovery for a dynamic fleet and alerting into Slack. You'll drive SLOs, on-call hygiene, and getting to root cause fast when a sensor in a facility misbehaves.
Security. Secrets and credentials in HashiCorp Vault, network access over Tailscale, and sensible least-privilege practices across cloud and edge.
MLOps plumbing (in partnership with the ML team). MLflow model registry and artifact tracking (models, configs, runs), and the provisioning/OTA path that gets a validated model onto the right sensors. You'll build and operate this infrastructure; the ML team owns model pipelines and the quality/automation gate. Together you'll push our deploy-time down from ~15 days toward same-day, plug-and-play deployment.
What success looks like (first 6–12 months)- Model/artifact tracking (MLflow) productionized and adopted as the single source of truth for models, configs, and runs.
- Fleet provisioning and OTA turned into a repeatable, observable, low-touch process as sensor count scales.
- Time-to-deploy a model to a sensor cut materially (our roadmap targets 5 days/sensor by end of 2026, trending toward under a day) — the infrastructure half of that is yours.
- No single point of failure on infra, DB, or deployment: everything documented, reproducible, and runnable by more than one person.
3–5 years of professional experience, including at least one project where you set up and operated a live, running production infrastructure end to end — plus hands-on exposure to building an ML model and/or active software development.
Strong, hands-on DevOps / platform engineering background — you've owned production cloud infrastructure (ideally AWS) as the person accountable for it.
Fullstack development in TypeScript — comfortable across a React/Next.js frontend and Node/Fastify (or similar) backend services with a relational database (Postgres/Supabase).
Working understanding of MLOps — model registries (MLflow), artifact tracking, and how models get deployed and served — enough to own the plumbing and collaborate closely with the ML team.
Infrastructure as code in anger — Terraform/OpenTofu, Pulumi, and/or Ansible — plus solid CI/CD pipeline design (GitHub Actions or equivalent).
Python packaging & distribution — building, versioning, and publishing installable Python packages (setuptools/pyproject.toml, internal package registries) as a deployment mechanism.
Containers for cloud services: Docker plus an orchestration/compute platform (ECS/Fargate, Kubernetes, or similar).
Observability with the Grafana/Prometheus ecosystem (or close equivalents) — metrics, logs, dashboards, alerting.
Comfortable operating relational databases in production (PostgreSQL preferred), including backups, migrations, and performance.
Linux fluency and comfortable scripting in Python and/or Bash.
Able to work independently and set the standard — this role owns the platform, it doesn't wait to be told how.
Nice-to-haveEdge / IoT / fleet experience: managing many remote devices, OTA updates, mesh networking (Tailscale/WireGuard).
Industrial / hardware exposure: PLCs (Siemens S7 / Snap7), MQTT (Mosquitto), device I/O, real-time constraints.
Deeper MLOps experience: hands-on model serving/provisioning at scale, Weights & Biases.
Frontend depth with Mantine, Chart.js/Plotly, or client-side PDF generation (react-pdf).
- TimescaleDB or other time-series databases; gRPC.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search