Distributed Systems Architect: Bare-Metal and High-Scale
Indexed description
What You’ll Do
Open-Source Infrastructure Mapping
- Map Analog’s Azure PaaS components to open-source equivalents: Event Hub to Kafka/Redpanda, IoT Hub to EMQX, ADX to ClickHouse or Apache Druid, and Blob Storage to Ceph/MinIO
- Produce side-by-side equivalence assessments documenting feature gaps, operational differences, and migration risk for each component transition
- Validate client hardware specifications and assess compliance with air-gap security requirements prior to each deployment
- Design multi-node clustering, rack-aware replication, and quorum topologies for bare-metal Kafka/Redpanda and PostgreSQL (Patroni) clusters
- Design resilience protocols for hardware and network switch failures in physically isolated or air-gapped client environments
- Architect RocksDB state backend tuning for stateful Apache Flink workloads, including compaction strategy, block cache sizing, and write-ahead log configuration
- Tune Linux kernel parameters, NUMA bindings, network ring buffers, and NVMe I/O scheduler configurations to eliminate hardware bottlenecks under petabyte-scale write workloads
- Define CPU affinity, IRQ balancing, and huge page configurations for latency-sensitive broker and database processes
- Author reproducible benchmark harnesses to validate configuration changes against client hardware before production rollout
- Deliver precise, implementation-ready configuration blueprints and runbooks for Terraform and Ansible automation, documents the DevOps team can execute without architectural interpretation
- Maintain a library of parameterized reference architectures covering single-rack, multi-rack, and geographically distributed bare-metal topologies
- Conduct pre-deployment reviews of client hardware specs and provide go/no-go assessments with remediation guidance
- 10+ years architecting and operating distributed infrastructure at TB to PB scale in production environments
- Expert-level Kafka or Redpanda clustering: partitioning strategy, ISR tuning, consumer group management, compaction policies, and operational lifecycle at scale
- Deep Linux systems engineering: kernel networking subsystems (TCP buffer tuning, interrupt coalescing), storage fabrics, and NUMA-aware process binding
- Proven track record deploying HA database clusters without cloud load balancers: Patroni, Pacemaker, or equivalent in production
- Experience deploying software-defined storage (Ceph, MinIO) on physical hardware in production, including CRUSH maps, erasure coding, and performance tuning
- Ability to produce implementation-ready blueprints and runbooks; your output must be directly actionable by a DevOps automation team without architectural interpretation
This is a full-time, on-site role.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search