Back to search
DDN Linkedin · Posted 6d ago

Lead Software Engineer - Golang

India

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Overview

DDN India is seeking great candidates to join our dynamic team of passionate customer-enabling technologists!

This is an incredible opportunity to be part of a company that has been at the forefront of AI and high-performance data storage innovation for over two decades. DDN Storage is a global market leader renowned for powering many of the world's most demanding AI data centers, in industries ranging from life sciences and healthcare to financial services, autonomous cars, Government, academia, research and manufacturing.

DDN Storage is the global leader in AI and multi-cloud data management at scale. Our cutting-edge storage and data management solutions are designed to accelerate AI workloads, enabling organizations to extract maximum value from their data. With a proven track record of performance, reliability, and scalability, DDN Storage empowers businesses to tackle the most challenging AI and data-intensive workloads with confidence.

Our commitment to innovation, customer success, and market leadership makes this an exciting and rewarding role for a driven professional looking to make a lasting impact in the world of AI and data storage.

Job Summary

About the Role

We are looking for a Staff Engineer to serve as the technical architect and lead for DDN Horizon — our provider-agnostic AI consumption interface and control plane for GPU compute and storage, built on the MOSS platform. You will own the end-to-end technical design of a distributed, event-driven control plane spanning multiple Go services (horizon-ctrl, provider-engine, and the MOSS core), a NATS-based messaging fabric, and Kubernetes-native deployment across on-premises, air-gapped, and SaaS topologies.

This is a senior technical-leadership role. You will set architectural direction, define the service and API contracts other teams build against, and make the hard calls on decomposition, consistency, and provider abstraction. You will still be deep in the design and code — reviewing critical paths, prototyping the hard parts, and mentoring engineers — but your leverage comes from raising the technical bar across the program rather than owning a single component. You will partner closely with product, DevOps, security, and the MOSS platform team to keep Horizon coherent as it scales from Phase 1 into a multi-service platform.

Key Responsibilities

  • Own the overall technical architecture of Horizon Control across horizon-ctrl, provider-engine, and their integration with the MOSS platform (identity, tenancy, authorization, governance, audit, API gateway), balancing "start simple, decompose later" against long-term extensibility.
  • Define and evolve the provider-agnostic REST API surface (OpenAPI 3.0.3, code-generated clients and stubs) so that no backend-specific concept leaks to clients and backends can be swapped without client-facing changes.
  • Set the standard for the event-driven, asynchronous architecture over NATS — Core NATS Request-Reply, Core NATS pub-sub notifications, and JetStream durable streams — including subject namespace design, durable consumers, queue groups, idempotency, and exactly-once-effect semantics.
  • Drive the design of async operation execution: the operation queue, competing consumers, reconciliation loops, drift self-heal, health checks, and SSE event streams to clients.
  • Own the provider abstraction and id-mapping model that keeps orchestration-backend concepts (Rafay in Phase 1) fully contained, ensuring Horizon ids never leak backend ids and the backend can be replaced by touching a single service.
  • Establish authentication and authorization patterns across services — JWT validation via JWKS fetched over NATS (ES256/asymmetric keys), OAuth 2.0 (including Device Authorization Flow for CLI), RBAC, and IDP federation — with zero synchronous dependencies on the auth hot path.
  • Define the data isolation and persistence strategy — per-service key/value and Postgres stores, backend-agnostic store interfaces, and the migration path toward Infinia KV — so services never write to each other's stores.
  • Set encryption and secrets-management standards: TLS/mTLS between services, encryption at rest (Kubernetes Secrets with encryption configuration, key rotation), and FIPS 140-3-aligned cryptography across all deployment models.
  • Own the Kubernetes and container strategy: Helm-packaged Go services, horizontal scaling profiles, ingress/API-gateway routing, and deployment across on-prem, air-gapped, and hosted environments (including OKE on OCI for internal dev/QA).
  • Lead architecture and design reviews, define engineering standards and RFC processes, and mentor senior and staff engineers across the Horizon program.
  • Partner with product to translate PRD scope into phased technical roadmaps, and with security/DevOps to keep security and observability embedded in the delivery pipeline.

Required Skills & Experience

  • 12+ years building production distributed systems and backend services, with a track record of owning architecture for a multi-service platform at staff/principal level.
  • Deep expertise in Go (or a comparable systems language) for building control planes, API services, and concurrent async workers.
  • Strong REST API design skills, including OpenAPI/spec-first workflows, code generation, versioning, and building provider-agnostic surfaces that hide backend detail.
  • Hands-on experience designing event-driven and message-driven systems with a broker such as NATS/JetStream, Kafka, or RabbitMQ — request-reply, pub-sub, durable queues, competing consumers, and idempotent/at-least-once delivery semantics.
  • Expert-level Kubernetes and container knowledge — Helm, ingress/gateway, scaling, health/readiness, secrets, and operating workloads across cloud, on-prem, and air-gapped environments.
  • Solid grounding in authentication and authorization at scale — OAuth 2.0, JWT/JWKS, RBAC, mTLS, and multi-tenant identity/tenancy models.
  • Practical experience with applied cryptography and encryption — TLS/mTLS, encryption at rest, key management and rotation, and FIPS-aligned or otherwise regulated cryptographic requirements.
  • Experience designing for multi-tenancy — tenant isolation, data partitioning, and per-tenant policy/quota enforcement.
  • Demonstrated ability to lead through technical influence: authoring designs and RFCs, running architecture reviews, and mentoring senior engineers.
  • Excellent written and verbal communication — able to make complex distributed-systems trade-offs legible to engineers, product, and executives.

Nice to Have

  • Experience with AI/ML infrastructure — GPU compute provisioning, workload orchestration (Kubernetes, SLURM), or high-performance storage.
  • Familiarity with orchestration/provisioning backends such as Rafay, or with building pluggable provider/adapter architectures.
  • Experience with reconciliation-based / declarative control planes (Kubernetes operators, controller patterns) and drift detection.
  • Background in on-prem, sovereign, government, or air-gapped delivery models.
  • Experience with S3-compatible object storage, key/value stores, or Postgres at scale, and with backend-agnostic storage abstractions.
  • Familiarity with OCI (OKE, DRG/IPSec VPN, Vault, OCIR) or comparable cloud networking and deployment.
  • Experience with observability and audit pipelines — structured logging (Fluent Bit), metrics, CloudEvents-style audit streams, and SSE.
  • Contributions to open-source distributed-systems, Kubernetes, or messaging projects.


Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search