Linux Platform & Observability Team Lead
Indexed description
At Eclit, we offer managed IT services that meet the end-to-end IT needs of businesses of all sizes with our self-developed technology platforms through Turkey's largest and most reliable cloud infrastructure.
With more than 20 years of experience, we are an industry-leading technology company. We simplify the complex structure of IT with our vision of setting the standards of the managed services industry in Turkey, our expert human resources and technological infrastructure.
Join us to be a part of this dynamic team and share our goal of taking technology to the next level with creative solutions!
What You Will Do
- Provide technical leadership to the Linux Platform & Observability team, ensuring reliable, scalable, and sustainable platform operations
- Manage the performance, security, high availability, and operational continuity of Linux infrastructure
- Support the technical growth of team members through mentoring, coaching, performance feedback, goal setting, and career development
- Design, develop, and continuously improve monitoring and observability platforms
- Lead the implementation, enhancement, and standardization of monitoring platforms, including Zabbix, Prometheus, Grafana, and related technologies
- Manage metrics collection, alert optimization, dashboard design, and capacity planning across the monitoring infrastructure
- Develop Ansible-based automation solutions to reduce operational workload and improve standardization
- Implement event correlation, noise reduction, and intelligent alerting capabilities
- Leverage AI-powered operational technologies to enhance incident detection, root cause analysis (RCA), predictive analytics, and self-healing capabilities
- Generate actionable insights from monitoring data to enable proactive operations
- Provide technical leadership during critical incidents, lead root cause analysis (RCA), and implement permanent improvements
- Monitor operational KPIs, SLAs, and service quality while driving continuous improvement initiatives
- Lead the creation and maintenance of technical standards, operational processes, and documentation
- Stay up to date with emerging Linux, monitoring, automation, and AI technologies and introduce relevant innovations to the team
- Take an active role in technical recruitment and contribute to building a strong engineering culture
Who You Are
- Bachelor's degree in Computer Engineering, Software Engineering, Electrical & Electronics Engineering, or a related field
- Minimum of 10 years of experience in Linux Platform, Systems Administration, or Infrastructure Operations, including at least 2 years of people management experience
- A people-focused leader with proven experience in mentoring, coaching, performance management, career development, and team growth
- Demonstrated success in building high-performing engineering teams and fostering a strong technical culture
- Strong hands-on experience with enterprise-scale Linux platforms, including high availability (HA), performance optimization, security, and capacity planning
- Strong experience in automation development, particularly with Bash/Shell scripting
- Hands-on experience with Ansible for Infrastructure as Code (IaC), Configuration Management, and operational automation
- Experience with virtualization and container technologies, including KVM, Docker, and Kubernetes
- Knowledge of cloud-native architectures and modern Linux platforms
- Advanced experience with observability and monitoring platforms, including Zabbix, Prometheus, Grafana, Alertmanager, and VictoriaMetrics/Thanos (preferred)
- Experience designing monitoring architectures with a solid understanding of metrics, logs, and distributed tracing
- Knowledge of or hands-on experience with AIOps, intelligent monitoring, event intelligence, predictive monitoring, and AI-powered operations
- Experience applying, or a strong interest in applying, AI-driven operational solutions using technologies such as OpenAI, LLMs, MCP, Agentic AI, Generative AI, and Microsoft Copilot to improve operational efficiency
- Experience developing monitoring automation, self-healing, event correlation, and root cause analysis (RCA) solutions is considered a strong advantage
- Solid understanding of incident, problem, change, major incident, and SLA management processes in line with ITIL best practices
- Strong analytical thinking, technical problem-solving, and prioritization skills
- Excellent communication and stakeholder management skills, with the ability to collaborate effectively across technical teams and business units
- Customer-oriented, results-driven, and committed to a culture of continuous improvement
- Good command of written and spoken English
What We Offer
- Agile and Dynamic Work Environment: A flexible, innovative work culture that nurtures your creativity and energy
- Continuous Development Opportunities: We support your career with free online courses and in-house development programs
- Extra Leave Benefits for Ecliters: At Eclit, we value important moments in your life by offering special leave beyond legal entitlements
- Comprehensive Health Insurance
- Gifts for Special Occasions: Including marriage and childbirth
- Flexible Benefits: Personalize your work-life balance with benefits tailored to your needs
- Monthly Meal Allowance: You can choose to receive your meal allowance as a meal ticket or as an additional payment to your salary
- Games and Rewards, Parties and Events: Strengthen team spirit with fun activities, prize competitions, and social events, where we work hard and play hard
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search