Site Reliability Engineering (SRE)
Indexed description
Core Responsibilities
1. Platform onboarding (first phase – critical)
- Control plane architecture
- Agentic workflows & orchestration
- Observability stack (logs, metrics, traces)
- Deployment pipelines & environments
- Build operational knowledge of real use cases, not just infra
2. Run & Support (steady state)
L1 / L2 ownership:
- Incident triage & resolution
- Monitoring platform health (SLA, latency, errors)
- Managing alerts & escalation flows
- Basic remediation (restart services, config fixes, rollback)
Operational excellence:
- Improve runbooks
- Reduce MTTR
- Identify recurring issues (problem management)
3. L3 Interface with the team
- Escalate complex issues (design flaws, bugs, scaling limits)
- Provide structured feedback (logs, reproduction steps, impact)
- Act as bridge between Global IT and platform engineering
Required Skills
SRE / Platform Ops fundamentals
- Incident management (ITIL mindset)
- Observability tools (Datadog, Prometheus, Grafana, etc.)
- Cloud environments (GCP/AWS/Azure)
- CI/CD understanding
- Strong plus
API-based systems & distributed architectures
- Event-driven systems / microservices
- Understanding of AI/LLM-based systems (at least operationally)
- Strong kubernetes knowledge, including cluster management an scaling
- Background in infrastructure as code
- Knowledge of GitOps
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search