Back to search
Duplo Linkedin · Posted 5d ago

Senior Site Reliability Engineer

Lagos

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Duplo is a Lagos-based fintech startup that enables businesses in Africa to automate their spend management, simplify cross-border payments, and control business finances all on one platform.


We want to make B2B payments as simple as P2P payment apps. Most business payments in Africa are made offline…yikes. We are on a mission to transform this. We are backed by top investors, including Tribe Capital, Commerce Ventures, Liquid2 Ventures, My Asia VC, Soma Capital, Y Combinator, Oui Capital, and others.


This is a unique opportunity. You'll have the responsibility and resources to take a significant part in the creation of a paradigm-changing product that will impact millions.


Key Responsibilities:


  • Lead technical design and implementation for large, high-risk, or high-complexity initiatives that improve the availability, scalability, and resilience of our backend systems
  • Design robust service boundaries, define maintainable infrastructure architecture, and guide reliability standards and best practices across the engineering team
  • Understand edge cases in money movement and design with operational safety and compliance in mind, not just uptime
  • Own observability end-to-end metrics, logging, tracing, and alerting so issues are caught before customers feel them
  • Build and refine incident response processes, lead incident diagnosis and resolution in production, and drive thorough postmortems that prevent recurrence
  • Work seamlessly with cross-functional partners including product, compliance, finance operations, support, backend engineering, and external providers, communicating clearly and adapting your approach to whoever you are working with
  • Collaborate closely with fellow engineers through code and infrastructure review, pairing, and technical discussion, giving and receiving feedback in a way that raises the quality of the team's work, not just your own
  • Make sound, well-reasoned decisions on architectural trade-offs, deployment strategies, and rollout safety, and know when to escalate decisions with broader business impact
  • Drive capacity planning, performance tuning, and cost-efficiency across infrastructure, balancing reliability against operational spend
  • Define and track SLOs/SLIs, and use them to prioritize reliability work against feature delivery
  • Mentor other engineers; raise engineering and operational standards and strengthen both the systems you own and the people you work with


Requirements:


  • Minimum of 7 years of experience in SRE, DevOps, or backend infrastructure roles, with deep expertise in Kubernetes, Docker, and AWS in production environments
  • Strong scripting/programming ability (TypeScript and/or another language used for tooling, automation, or services) to build internal tools and automate operational work
  • Practical expertise with PostgreSQL operations at scale, including replication, backup/recovery, performance tuning, and patterns that protect correctness in financial systems
  • Practical expertise with Redis for caching, coordination, rate-limiting, and distributed locking, including operating it reliably in production
  • Strong command of observability tooling (e.g., Prometheus, Grafana, Datadog, ELK, or equivalent) and a track record of building monitoring that catches problems early
  • Demonstrated experience designing and maintaining CI/CD pipelines, infrastructure as code (e.g., Terraform), and secrets management
  • Demonstrated experience in fintech or other high-stakes systems, including provider integration resilience, failover design, and compliance-sensitive infrastructure
  • A track record of owning complex, high-stakes incidents from detection through resolution and postmortem, and standing behind the outcomes
  • Strong communication skills, with the ability to work effectively across engineering, product, compliance, and external partners
  • A demonstrated habit of catching issues before they become incidents, and of treating failure as something to learn from rather than deflect
  • Experience mentoring or raising the operational standard of a team, not just delivering individually

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search