Site Reliability Engineer
Indexed description
You will bridge the gap between development and operations by applying a software engineering mindset to system administration.
Working with modern container and cloud technologies, while shaping the development landscape of the future.
Main Responsibilities
- Design, implement, and refine advanced monitoring dashboards and alerting systems (using Azure App Insights, Prometheus, Grafana, or Datadog) to ensure 24/7 visibility
- Act as the first line of defense for production issues. You will lead troubleshooting efforts, perform Root Cause Analysis (RCA), and document incidents to prevent recurrence
- Manage technical escalations and tickets. You will communicate directly with clients to resolve complex technical issues, acting as a trusted technical advisor
- Support the deployment of builds through CI/CD pipelines, ensuring that every release meets our stability and performance standards
- Identify manual operational tasks and automate them using scripting (Bash, PowerShell) or Infrastructure as Code (Terraform, Ansible)
- Monitor KPIs, manage database health (KQL, MSSQL), and optimize containerized environments (Docker/Kubernetes)
- 3+ years in an SRE, DevOps, or high-level Systems Operations role
- Degree or Diploma in Computer Science or a related field (MSc preferred)
- Proven experience managing production workloads in Public Cloud environments (Azure, AWS, or GCP). Azure experience is highly preferred
- Extensive experience with metrics and logging systems (Azure Monitor, Logstash, Dynatrace, or similar)
- Hands-on experience with Azure (preferred) and container orchestration (Kubernetes/Docker)
- Strong scripting skills (Bash, or PowerShell)
- Familiarity with .NET/ASP.NET environments is a strong plus
- Knowledge of Infrastructure as Code (Terraform, Ansible, or Chef)
- Comfort with KQL and relational databases (MSSQL/PostgreSQL)
- Excellent command of English and the ability to explain complex technical issues to both developers and customers
- Organizational
- Time management
- Analytical thinking
- Work under pressure
- Good interpersonal relationships
- Complete tasks with minimal supervision
- Competitive salary and Private Medical-Life Insurance
- An individual and well-structured introduction and training when you onboard
- Training budget for personal & professional development
- Friendly and highly motivated colleagues
- Hybrid working model
- Career progression opportunities
- Continuous competencies development
- A spacious workplace that promotes co-creation, collaboration, and improvement.
- Flexible working hours
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search