Principal Platform Engineer
Indexed description
Principal Platform Engineer (Distributed Systems & Database Infrastructure)
Hybrid - London
Salary - £120,000 - £140,000 + Bonus
We are seeking an experienced Principal Platform Engineer to join a highly skilled infrastructure engineering team responsible for the design, deployment and operation of mission-critical distributed platforms that power some of the organisation's most important services.
This is a hands-on engineering role focused on distributed systems, cloud infrastructure, platform engineering and large-scale data platforms. The successful candidate will be a technical leader with deep expertise in AWS, Terraform and distributed persistence technologies, capable of influencing architecture decisions while remaining close to implementation and operational ownership.
You will play a key role in building resilient, highly available and scalable platforms supporting critical business workloads across both cloud and hybrid environments.
The Opportunity
This role sits within a specialist engineering function responsible for designing and operating distributed platforms at scale.
You will work closely with software engineers, architects, SRE teams and business stakeholders to modernise infrastructure, improve platform reliability and deliver robust cloud-native solutions.
The ideal candidate will have significant experience with distributed databases, event-driven systems and highly available infrastructure environments alongside strong AWS and Infrastructure-as-Code expertise.
Key Responsibilities
- Design, build and operate large-scale distributed infrastructure platforms.
- Develop and maintain highly available cloud-native environments using AWS and Terraform.
- Deploy, manage and support distributed persistence technologies and data platforms.
- Design resilient architectures incorporating clustering, replication, failover and disaster recovery patterns.
- Improve platform reliability, scalability, observability and operational excellence.
- Troubleshoot complex infrastructure and distributed systems issues across production environments.
- Drive automation and Infrastructure-as-Code adoption across engineering teams.
- Define and implement platform engineering best practices and engineering standards.
- Mentor engineers and act as a subject matter expert for distributed systems and cloud infrastructure.
- Collaborate with software engineering teams to ensure platforms are scalable, secure and fit for purpose.
Technology Environment
- AWS
- Terraform
- Linux
- Kubernetes
- Docker
- Kafka / AWS MSK
- Cassandra
- Couchbase
- ScyllaDB
- MongoDB
- PostgreSQL
- Redis
- ClickHouse
- GitHub Actions / GitLab CI
- Prometheus
- Grafana
- OpenTelemetry
- Python
- Golang
What We're Looking For
Essential
- Strong experience building and operating AWS-based infrastructure platforms.
- Expert-level Terraform experience within enterprise production environments.
- Extensive Linux systems administration and troubleshooting capability.
- Deep understanding of distributed systems architecture.
- Hands-on experience designing, deploying and supporting distributed persistence platforms.
- Strong knowledge of clustering, replication, high availability and disaster recovery architectures.
- Experience operating critical production services at scale.
- Strong automation and Infrastructure-as-Code mindset.
- Excellent problem-solving and stakeholder engagement skills.
Highly Desirable
- Experience administering Cassandra, Couchbase, ScyllaDB, MongoDB or similar distributed databases.
- Experience with Kafka or other distributed messaging platforms.
- Understanding of CAP Theorem, consistency models and distributed database design principles.
- Platform Engineering and SRE experience.
- Experience supporting mission-critical financial services or transaction-processing systems.
- Exposure to multi-region or active-active architectures.
Experience within Financial Services, FinTech, Payments, Trading, Retail Technology or other large-scale distributed environments would be highly advantageous.
What Success Looks Like
- Highly resilient and scalable infrastructure platforms.
- Reliable distributed database and messaging services.
- Increased platform automation and operational efficiency.
- Improved service availability and recovery capabilities.
- Reduced operational risk through engineering best practices.
- Strong technical leadership and engineering mentorship.
- Delivery of secure, modern and cloud-native platform capabilities.
This is an excellent opportunity for a senior engineering professional with deep expertise in distributed systems, AWS, Terraform and distributed database technologies to influence the future direction of a highly scalable platform environment while remaining hands-on with technology.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search