Reliability Engineer
Indexed description
Remote-first (Europe – CET ±1 hour) | Full-time | Senior
About PlyoPlyo is a proptech company on a mission to revolutionize the real estate industry with immersive digital solutions that inspire, engage, and streamline property sales. We build tailored, all-in-one tools for property developers and real estate agencies—spanning 3D visuals, interactive project websites, and next-gen listing portals.
Our vision is to create the most immersive home-buying experience in Europe, blending photorealistic design with intuitive exploration tools. From single-project microsites to full-scale portals, our platforms prioritize authenticity, performance, and buyer engagement.
We are a remote-first product development team distributed across Europe, working in cross-functional squads that emphasize autonomy, clarity, and continuous delivery. To maintain healthy collaboration, we hire primarily within European time zones close to Oslo (CET ±1), and our working hours are anchored to Oslo time.
About the RoleAs our Reliability Engineer, you'll own production operations and delivery infrastructure for the Plyo platform — the multi-tenant real-estate SaaS behind our developer console and embeddable buyer SDK. This is a foundational hire: you'll take a modern, GitOps-shaped estate (GKE, ArgoCD, Helm, Terraform, OpenTelemetry) from "built by brilliant engineers on the side" to "operated as a product," just as our first Enterprise customers go live. You'll work directly with the CTO and Lead Engineer, and everything you build — dashboards, alerts, runbooks, provisioning — becomes the standard the whole platform team runs on.
What You'll Do- Own production reliability end to end: SLOs, alerting, incident response, and blameless postmortems for ~16 microservices on GKE across staging and production.
- Build production observability on our existing OpenTelemetry foundation — dashboards-as-code, RED metrics, frontend telemetry, and load-balancer-level visibility.
- Own our GitOps delivery machinery: Helm + ArgoCD, a gated three-step production promotion flow, and the Terraform estate on GCP (GKE, Cloud SQL/Postgres, Redis, Pub/Sub, KMS, BigQuery, CDN).
- Steward CI/CD for a 22-workspace TypeScript monorepo on GitHub Actions with self-hosted runners — pipeline health, speed, and cost.
- Keep environments trustworthy: a staging that's green by policy, and a 30-container local development stack the whole team boots daily.
- Design and run our on-call: Oslo-anchored, small-team-sane (severity-gated paging, alert-quality budgets — every page actionable).
- Turn operational toil into product: convert manual backfill/provisioning scripts into audited operator tooling, and own secret management and key rotation.
- Run upgrades as a program, not a fire drill: version policy, maintenance windows, EOL tracking, and the evidence trail Enterprise security reviews ask for.
Must-Have:
- 5+ years running production Kubernetes (GKE preferred) with GitOps (ArgoCD/Flux) and Helm
- Strong Terraform on GCP — you treat console-made resources as bugs
- Deep GitHub Actions experience at monorepo scale, ideally with self-hosted runners
- Observability fluency: OpenTelemetry, SLOs, Cloud Monitoring or equivalent — you've built alerting from zero before
- Enough TypeScript/Node to read the services you operate and contribute tooling
- Security-operations hygiene: secrets, workload identity, key rotation, incident handling
- An automation reflex — you measure toil and delete it
- Fluent English (written and spoken)
Nice-to-Have:
- You've operated platforms where AI agents are first-class engineers (agent-driven PRs, policy-gated automation) and you design guardrails, not gatekeeping
- Postgres operations at depth (multi-tenant architectures), Pub/Sub or event-driven systems, CDN/edge and signed-URL delivery
- FinOps / cloud cost management experience
- You've taken a platform through SOC 2 / ISO 27001 and know which controls are real engineering work
- You've been the person who ended a bus-factor-of-one situation somewhere else
- A product-driven, engineering-led culture
- Remote-first flexibility with async-friendly practices
- A foundational role with real ownership — you define how this platform is operated
- Stable teams and long-term product ownership
- Competitive region-based salary
- Annual learning budget
- Occasional team meetups and cross-team collaboration
- Intro call with our CTO
- Async technical task (~3 hours)
- Panel interview with our engineers
- Offer & onboarding 🎉
If you love making production boring, turning incidents into architecture, and building the reliability foundation a growing platform deserves — we'd love to hear from you.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search