Senior Site Reliability Engineer (MAAS)
Indexed description
Start date: ASAP
Languages: Fluent English required
Industry: Cloud Computing / GPU Infrastructure
About The Opportunity
Pragmatike is hiring a Senior SRE / Infrastructure Engineer to help operate and scale a distributed infrastructure platform spanning bare-metal GPU nodes, Kubernetes, virtualization, networking, and multi-site environments.
You’ll work close to the infrastructure itself, from BMCs, hardware and MAAS provisioning through Kubernetes, networking, observability, automation, and site operations.
This is a hands-on role for someone who enjoys owning infrastructure end-to-end, building reliable systems, and automating everything that can be automated. Startup or hyper-growth experience is a strong plus: autonomy, ownership, and speed matter here.
What You’ll Do
- Operate and maintain large-scale Linux infrastructure across Debian/Ubuntu-based bare-metal and virtualized environments.
- Own MAAS-based bare-metal provisioning, including region/rack controllers, PXE, commissioning, cloud-init, node lifecycle, and API/CLI automation.
- Operate and maintain production Kubernetes clusters, including upgrades, node pools, networking, storage, security hardening, and troubleshooting.
- Design and maintain multi-site networking across VLANs, L2/L3 routing, bonded interfaces, VPNs, firewalls, and DNS.
- Automate infrastructure provisioning and operations using Ansible, Bash/Python, OpenTofu/Terraform, and Git-based workflows.
- Build and maintain automated deployment workflows including PXE, Preseed, and cloud-init.
- Operate observability platforms using Prometheus, Grafana, Alertmanager, VictoriaMetrics/VictoriaLogs, or comparable tooling.
- Define and improve SLIs, SLOs, alerting, and reliability practices across infrastructure and platform services.
- Lead infrastructure incident response, troubleshooting, escalation, and post-incident improvements.
- Maintain on-call processes and operational coverage across distributed environments.
- Work close to the hardware layer, including IPMI/Redfish, BMCs, RAID, storage, hardware diagnostics, and GPU infrastructure.
- Manage virtualization platforms including Proxmox, KVM/libvirt, OpenStack, or VMware, including GPU passthrough where required.
- Build and maintain internal infrastructure tooling for host discovery, configuration, IPAM, hardware health, and operational automation.
- Own infrastructure lifecycle activities including site onboarding, maintenance, decommissioning, drift detection, and operational runbooks.
- Work closely with engineering and cross-functional teams to improve reliability, resource utilization, and operational efficiency.
- 5+ years of hands-on SRE, Infrastructure, Systems, or Platform Engineering experience.
- Expert-level Linux administration, particularly Debian/Ubuntu.
- Strong production experience with MAAS and bare-metal provisioning.
- Expert-level, hands-on experience operating Kubernetes in production, including cluster lifecycle, networking, storage, upgrades, and troubleshooting.
- Strong network engineering skills across VLANs, L2/L3 routing, bonding, VPNs, firewalls, and DNS.
- Strong automation skills with Ansible, Bash and/or Python.
- Experience with Terraform/OpenTofu and Git-based infrastructure workflows.
- Production experience with Prometheus/Grafana or comparable observability platforms.
- Experience with incident response, on-call operations, monitoring, alerting, and reliability practices.
- Experience with Proxmox, KVM/libvirt, OpenStack, VMware, or comparable virtualization technologies.
- Experience with bare-metal hardware, BMCs, IPMI/Redfish, storage, and hardware troubleshooting.
- Strong understanding of distributed systems, container orchestration, and infrastructure reliability.
- Experience with infrastructure security including RBAC, firewalls, network policies, secrets management, and security hardening.
- Ability to create SOPs, runbooks, and operational processes from scratch.
- Comfortable working autonomously in a fast-paced, engineering-driven environment.
- Experience operating GPU infrastructure or GPU-heavy Kubernetes/bare-metal environments.
- Proxmox VE with ZFS/Ceph and GPU passthrough.
- VictoriaMetrics / VictoriaLogs or similar large-scale observability platforms.
- NetBox or other IPAM / infrastructure inventory platforms.
- Vault, SOPS, Atlantis, or similar infrastructure security and automation tooling.
- Experience with Cloudflare APIs, DNS automation, or tunnels.
- Experience with UniFi or comparable site networking platforms.
- Experience with service mesh or advanced CNI implementations.
- Go experience, particularly for infrastructure tooling or custom exporters.
- Experience with Ceph or distributed databases running on bare metal.
- Experience operating infrastructure across multiple sites and timezones.
- Experience establishing SRE frameworks and reliability practices in growing organizations.
- 100% remote with flexible working hours
- High-impact role with significant technical ownership and autonomy
- Work directly with bare-metal, Kubernetes, networking, and GPU infrastructure
- International, engineering-driven team
- Strong focus on automation, reliability, and infrastructure at scale
- Opportunity to shape the architecture and operational foundations of a growing cloud platform
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search