Monitoring Specialist
Indexed description
We’re driven by speed, ambition, and bold ideas. At our core, we’re product creators and problem-solvers who thrive in a high-performance environment. We’re looking for exceptional talent—smart, adaptable, and motivated individuals eager to make an impact, scale business functions at pace, and continuously improve.
The Role
Keep our global gaming platform running smoothly by monitoring systems, diagnosing issues, and ensuring seamless performance for millions of players worldwide.
What You Will Be Doing
- Monitor infrastructure and application health using tools like Prometheus, Grafana, and across our multi-site production environment
- Analyze metrics, logs, and alerts to detect and resolve issues before they impact our players
- Perform Root Cause Analysis on incidents and document solutions for future reference
- Optimize monitoring systems by fine-tuning metrics and reducing alert noise
- Collaborate with development and operations teams during incident response
- Participate in incident reviews and contribute to continuous improvement initiatives
- Solid technical foundation in system administration, DevOps, or technical support roles
- Strong understanding of server, network, and application performance metrics (CPU, memory, latency, RPS)
- Experience with log analysis tools such as Elasticsearch, Kibana, Loki, or Splunk
- Hands-on experience configuring monitoring tools like Prometheus, Alertmanager, or Datadog
- Proficiency with Linux systems and command-line operations
- Flexibility to work shift-based schedules during non-business hours (European time alignment)
- Willingness to work in a shift-based schedule
- Kubernetes hand-on experience
Benefits:
- Full remote work flexibility
- Generous paid time off including vacation, sick leave
- Engaging company events and vibrant culture
- Unlimited growth opportunities in a fast-scaling business
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search