Staff Software Engineer, Platform Metering
Indexed description
This role sits at the intersection of cloud infrastructure, distributed systems, billing, and platform reliability. You’ll own the domain-level architecture for metering provisioned resources such as GPU compute, Kubernetes nodes, bare-metal capacity, storage, and future platform services. Your work will ensure that every billable resource has a trustworthy usage trail: accurate enough for billing, timely enough for credit enforcement, and explainable enough for customers, finance, and engineering teams. This is an opportunity to define a foundational platform domain early, setting the metering architecture and standards that future Nscale services will build on.
Cloud metering powers Nscale’s usage-based billing platform by producing accurate, deduplicated, auditable usage records for rating, credit burn-down, entitlements, cost attribution, margin analysis, and customer-facing usage dashboards.
What you’ll work on
- GPU and compute metering: measuring provisioned GPU-hours across instances, Kubernetes, Slurm, bare metal, and future compute products.
- Storage metering: measuring GiB-hours for file storage, object storage, and future storage products.
- Usage event production: building controllers and resource watchers that emit durable service consumption events into the metering pipeline.
- Usage attribution: mapping consumption to organizations, projects, regions, clusters, flavors, allocations, and workloads.
- Credit and customer safeguards: providing the usage signals needed for prepaid credit burn-down, headroom checks, grace periods, customer notifications, and fair enforcement when credit is exhausted.
- Reconciliation and visibility: ensuring usage can be audited, replayed, explained, and surfaced to customers, finance, support, and engineering.
- Domain-level technical direction. Set the technical direction for platform metering across a defined domain, influencing engineering squads across Nscale that produce or consume usage data.
- Design accurate metering models. Define how Nscale measures resources over time, including instance-hours, GPU-hours, GiB-hours, readiness states, allocation burn-down, and future shared-pool or pod-level usage models.
- Build reliable event producers. Design and implement controllers and resource watchers that emit self-contained usage quanta for compute, Kubernetes, bare metal, storage, and other platform resources.
- Engineer for idempotency and auditability. Ensure metering events can survive retries, redelivery, backfills, partial outages, and customer disputes through deterministic transaction IDs, provenance, traceability, and reconciliation.
- Integrate with billing systems. Partner with billing, product, finance, and platform teams to ensure usage events can be rated, aggregated, credited, and surfaced to customers and internal stakeholders.
- Create leverage through standards. Establish shared schemas, libraries, conventions, dashboards, and operational runbooks so that new platform services can add metering consistently.
- Own production outcomes. Operate the metering platform with strong observability, alerting, incident response, data-quality checks, and reconciliation against billing records and infrastructure state.
- Mentor and influence. Guide other engineers through architecture reviews, implementation choices, and operational best practices for high-integrity usage systems.
- Extensive experience designing, building, and operating distributed systems in production, ideally in cloud infrastructure, data platforms, billing, control planes, or platform engineering.
- Experience with resource lifecycle, capacity, or usage tracking systems such as compute instances, Kubernetes nodes, storage volumes, jobs, workloads, quotas, or entitlements.
- Strong understanding of event-driven architecture, including reliable delivery, idempotency, aggregation, replay, and failure handling.
- Strong operational discipline, including monitoring, alerting, incident response, reconciliation, data-quality checks, and post-incident improvement.
- Proven ability to lead ambiguous technical work across team boundaries and drive domain-level delivery through influence rather than formal authority.
- Strong software engineering fundamentals, with proficiency in typed backend or systems languages. Our primary stack is Go, with some services in Rust and Python.
- Comfortable working in a fast-paced, ambiguous environment with high ownership, pragmatic judgement, and a bias toward measurable business impact.
- You use AI tools like Claude or Cursor as a core part of your development workflow to create leverage, increase quality, and accelerate delivery.
- Experience with cloud billing, chargeback/showback, prepaid credit systems, entitlements, quota enforcement, or customer-facing usage dashboards.
- Experience with GPU cloud infrastructure, Kubernetes, bare-metal provisioning, workload scheduling, storage platforms, or AI/ML inference and training workloads.
- Experience with billing, usage, ledger-style, event-sourced, or replayable data systems that support auditability, reconciliation, backfills, and dispute investigation.
- Experience with Kubernetes controllers/operators, controller-runtime, CRDs, admission webhooks, or multi-cluster resource watchers.
- Experience with infrastructure-as-code, cloud providers, regional control planes, and service catalogs or flavor/rate-card models.
- Strong product sense for making usage data understandable to customers, finance, and support teams, especially when investigating billing disputes or consumption anomalies.
Salary Range: $220,000 USD - $266,667 USD
For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search