Back to search
HCLTech Germany Linkedin · Posted 23d ago

Site Reliability Engineer (f/m/d) | Grafana, ELK & Kubernetes | Remote Germany

Germany

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Join HCLTech Germany as a Site Reliability Engineer


At HCLTech, we're powering digital transformation for some of the world's largest enterprises. We are looking for a talented Site Reliability Engineer (SRE) who is passionate about reliability, automation, observability, and cloud-native operations.

If you enjoy solving complex production challenges, automating repetitive tasks, building scalable monitoring solutions, and ensuring business-critical services remain available 24/7, we'd love to hear from you.


Location: Germany (Remote)

Employment Type: Full-Time

Candidates must be based in Germany or willing to relocate to Germany.


What You'll Do:


As a Site Reliability Engineer, you will play a key role in ensuring the stability, availability, performance, and security of mission-critical platforms.


Your responsibilities will include:

  • Designing and improving monitoring and observability solutions using Grafana, Prometheus, and ELK Stack
  • Managing and optimizing Elasticsearch/OpenSearch clusters
  • Supporting and operating containerized workloads in Kubernetes environments
  • Developing automation solutions using Python, Go, or Bash
  • Managing CI/CD pipelines with tools such as Jenkins, Helm, and ArgoCD
  • Leading troubleshooting efforts across infrastructure, platform, and application layers
  • Participating in 24x7 incident response and on-call rotations
  • Performing root cause analysis and driving continuous service improvements
  • Creating and maintaining operational documentation and runbooks
  • Collaborating with global engineering, cloud, and operations teams


What We're Looking For:


Must-Have Skills:

  • Grafana & ELK Stack (mandatory)
  • Kubernetes administration and container platforms
  • Linux system administration
  • Elasticsearch and/or OpenSearch
  • Prometheus monitoring ecosystem
  • CI/CD tooling (Jenkins, Helm, ArgoCD)
  • Infrastructure automation and scripting
  • Python, Bash, or Go
  • Understanding of networking concepts and REST APIs
  • Strong troubleshooting and incident management experience


Preferred Qualifications:

  • AWS and/or Azure experience
  • Kubernetes certifications (CKA/CKAD)
  • Elastic Certified Engineer
  • LPIC Level 2 or equivalent Linux certifications
  • Experience working in enterprise-scale production environments


Important Requirement:

This role supports environments subject to German security regulations.

Applicants must be eligible for and willing to undergo the German Ü2 Security Clearance process.

EU + NATO citizenship is required.


Why HCLTech?

Be part of a global technology leader

  • 225,000+ employees across 60 countries
  • Work with enterprise-scale cloud and infrastructure platforms
  • Remote working model within Germany
  • International and diverse teams
  • Continuous learning and certification opportunities
  • Long-term career growth in cloud, platform engineering, and reliability engineering


At HCLTech, we believe technology is powered by people. We foster an inclusive workplace where innovation, collaboration, and personal growth drive success.

Apply now and help us build reliable, scalable, and secure platforms that power global businesses every day.

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search