Senior Platform Engineer (m/f/d)
Indexed description
We are seeking a hands-on Senior IT Platform Engineer to support and operate a Kafka-based event streaming platform. This role focuses on day-to-day operations, incident resolution, and platform reliability, ensuring seamless service for internal users. You will work closely with engineering teams, platform owners, and internal customers to maintain system stability, improve performance, and enhance self-service capabilities.
How you create impact
Platform Operations & Reliability
- Operate, monitor, and maintain a Kafka-based messaging platform hosted on AWS (EKS).
- Ensure high availability, performance, and resilience of the platform.
- Monitor system health using logs, metrics, and observability tools (e.g., Grafana, Loki).
- Support scaling of infrastructure to handle increasing workloads and integrations.
- Investigate and resolve platform incidents across Kafka components (brokers, producers, consumers).
- Analyze logs and metrics to identify root causes and implement fixes.
- Escalate complex issues to engineering teams when needed.
- Participate in on-call (L3 support) rotations.
- Execute operational tasks following runbooks and SOPs.
- Perform Kafka configuration tasks (topics, ACLs, schemas) using defined processes.
- Continuously improve operational procedures and documentation.
- Contribute to post-incident reviews and preventive improvements.
- Act as a primary contact for internal platform users.
- Support users via collaboration tools (Slack, Teams).
- Provide technical guidance on platform usage and best practices.
- Translate user issues into actionable technical requirements.
- Support CI/CD pipelines (GitLab) for infrastructure and configuration changes.
- Assist in automating infrastructure using GitOps (ArgoCD) and Helm.
- Improve platform observability, performance, and operational efficiency.
- Enable self-service capabilities for teams (authentication, access, topic management).
- Experience with Apache Kafka and its ecosystem (e.g., Kafka Connect, Schema Registry)
- Strong understanding of distributed systems (partitioning, replication, scaling, fault tolerance)
- Hands-on experience with Kubernetes and container orchestration (EKS preferred)
- Experience with cloud platforms, preferably AWS (VPC, EKS, S3, IAM, networking, private links)
- Familiarity with CI/CD pipelines (GitLab) and version control using Git
- Understanding of Infrastructure as Code (IaC) and GitOps practices (Terraform, Helm, ArgoCD)
- Proven experience in IT operations, platform support, or production support environments
- Strong troubleshooting skills using logs, metrics, and monitoring tools
- Experience with incident management processes, escalation, and root cause analysis
- Familiarity with runbooks, SOPs, and structured operational workflows
- Experience supporting high-availability, distributed production systems
- Experience with observability and monitoring tools such as:
- Grafana, Prometheus
- Loki, Mimir, Alloy (LGTM stack)
- Ability to analyze system behavior, performance, latency, and throughput
- Scripting skills in Bash, Python, or Go for automating operational tasks
- Experience automating infrastructure provisioning and configuration
- Experience managing:
- Kafka clusters, topics, ACLs
- Schema registries and data contracts
- Understanding of asynchronous communication patterns and event-driven architecture
- Basic understanding of the Java ecosystem
- Experience with container tooling and cloud-native platforms
- Exposure to data integration platforms and large-scale messaging systems
- Strong problem-solving and analytical thinking
- Excellent communication skills (written and spoken English)
- Customer-focused mindset with ability to support internal users effectively
- Ability to collaborate across teams and contribute to shared decision-making
- Proactive attitude toward continuous improvement and automation
- Willingness to learn and grow in a platform engineering environment
Who we are
Logistics shapes everyday life - from the goods we consume to the healthcare we rely on. At Kuehne+Nagel, your work goes beyond logistics; it enables both ordinary and special moments in the lives of people around the world.
As a global leader with a strong heritage and a vision to move the world forward, we offer a safe, stable environment where your career can make a real difference. Whether we help deliver life-saving medicines, develop sustainable transportation solutions or support our local communities, your career will contribute to more than you can imagine.
We kindly advise that placement agencies refrain from submitting unsolicited profiles. Any submissions of candidates without prior signed agreement will be considered our property and no fees will be paid.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search