Platform Engineer (m/f/d)
Indexed description
As a Platform Engineer (m/f/d), you will be responsible for supporting, operating, and maintaining a Kafka-based event streaming platform, ensuring its reliability, performance, and availability. You will manage day-to-day platform operations, resolve incidents, and work closely with engineering teams, platform owners, and internal users to improve platform stability, optimize performance, and enhance self-service capabilities.
How you create impact
- Operate, monitor, and maintain a Kafka-based messaging platform on AWS (EKS), ensuring high availability, performance, resilience, and scalability.
- Monitor platform health using observability tools such as Grafana and Loki, proactively identifying and addressing performance issues.
- Investigate, troubleshoot, and resolve incidents across Kafka components (brokers, producers, and consumers), escalating complex issues when required and participating in L3 on-call support.
- Execute operational tasks and Kafka configuration changes (topics, ACLs, schemas) by following established runbooks, SOPs, and best practices.
- Continuously improve operational processes, documentation, and post-incident reviews to enhance platform reliability and prevent recurring issues.
- Provide technical support and guidance to internal users through collaboration tools, translating business needs into actionable technical solutions.
- Support CI/CD pipelines using GitLab and contribute to infrastructure automation with GitOps (ArgoCD) and Helm.
- Drive platform improvements by enhancing observability, operational efficiency, and self-service capabilities, including authentication, access management, and topic provisioning.
- Hands-on experience with Apache Kafka and its ecosystem, including Kafka Connect, Schema Registry, and managing Kafka clusters, topics, ACLs, and data contracts.
- Strong understanding of distributed systems, event-driven architectures, and asynchronous messaging patterns, with experience supporting high-availability production environments.
- Experience with Kubernetes (EKS preferred) and AWS services, including VPC, IAM, S3, networking, and private connectivity.
- Familiarity with CI/CD pipelines (GitLab), Infrastructure as Code, and GitOps practices using Terraform, Helm, ArgoCD, and Git.
- Proficiency in monitoring and troubleshooting distributed systems using observability tools such as Grafana, Prometheus, Loki, Mimir, and Alloy, with strong incident management and root cause analysis skills.
- Scripting experience in Bash, Python, or Go to automate operational tasks, infrastructure provisioning, and platform configuration.
- Basic knowledge of the Java ecosystem, container technologies, cloud-native platforms, and large-scale data integration or messaging systems.
- Strong analytical, communication, and collaboration skills, with a customer-focused mindset, commitment to continuous improvement, and eagerness to learn and grow in a platform engineering environment.
Who we are
Logistics shapes everyday life - from the goods we consume to the healthcare we rely on. At Kuehne+Nagel, your work goes beyond logistics; it enables both ordinary and special moments in the lives of people around the world.
As a global leader with a strong heritage and a vision to move the world forward, we offer a safe, stable environment where your career can make a real difference. Whether we help deliver life-saving medicines, develop sustainable transportation solutions or support our local communities, your career will contribute to more than you can imagine.
We kindly advise that placement agencies refrain from submitting unsolicited profiles. Any submissions of candidates without prior signed agreement will be considered our property and no fees will be paid.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search