Senior Site Reliability Engineer
Indexed description
About Us
ColendiBank is Türkiye's first AI-powered digital native universal deposit bank, pioneering the integration of advanced technology with accessible financial services. Officially launched in 2024 following the acquisition of a coveted digital banking license, the bank caters to individuals and SMEs with a focus on micro credit and innovative deposit services. As a fully digital institution, ColendiBank leverages a cutting-edge, scalable core banking platform to deliver seamless, secure, and user-friendly experiences. Its commitment to innovation positions it as a leader in redefining digital banking in Türkiye.
Job Description
We are looking for a Senior Site Reliability Engineer to join our Platform & DevOps team. You will be responsible for the production reliability of ColendiBank's digital banking services, working closely with development, infrastructure, and security teams and our technology vendors. The role covers incident management, observability, capacity, Kubernetes operations, and automation.
Our environment:
- Kubernetes platforms and virtualized Linux infrastructure
- GitOps-based delivery and CI/CD pipelines
- Java-based microservices and core banking systems
- Relational databases, messaging, and caching systems
- APM, logging, and monitoring platforms
You do not need experience with every tool we use today. We expect you to bring practices and tools that have worked for you elsewhere and to help shape how our platform and reliability practices evolve.
Key Responsibilities
- Define SLIs and SLOs for critical customer journeys and track them together with DORA metrics to prioritize reliability work
- Lead major incident response, coordinating development, infrastructure, security, and vendor teams
- Own the incident management process: on-call, runbooks, and blameless postmortems with action items tracked to completion
- Improve observability across applications, infrastructure, and third-party platforms (logs, metrics, traces, synthetic checks) and keep alerting actionable
- Operate Kubernetes clusters (health, upgrades, resource management, autoscaling) and standardize deployments on GitOps
- Plan capacity and performance: resource and connection budgets, autoscaling behavior, and load testing before launches and campaigns
- Keep platform components up to date through planned upgrades and patches with tested rollback plans
- Automate repetitive operational work with infrastructure as code and configuration management
- Run disaster recovery and backup restore tests, and support regulatory and audit requirements
- Define operational readiness criteria that new services and third-party software must meet before going live
- Use AI tools to speed up incident analysis, root cause analysis, and automation work
What We're Looking For
- 5+ years in SRE, DevOps, or production engineering roles, preferably in banking, fintech, or another regulated environment
- Strong Linux administration and troubleshooting skills
- Hands-on experience operating production Kubernetes, with Helm, GitOps, and CI/CD pipelines
- Solid networking fundamentals: DNS, TLS, HTTP, load balancers, firewalls
- Experience troubleshooting Java and Spring Boot applications in production (memory, garbage collection, thread dumps, connection pools)
- Operational knowledge of relational databases (PostgreSQL or Oracle) and messaging or caching systems such as Kafka or Redis
- Experience with observability tools such as Dynatrace, ELK, Zabbix, or Grafana, and with designing actionable alerting
- Proven incident management experience (on-call, root cause analysis, postmortems) and willingness to join an on-call rotation
- Automation skills in Python, Go, or Bash; experience with Ansible or Terraform
- Understanding of security practices in production: least privilege, secrets management, patching
- Clear written communication and a strong sense of ownership
- Ability to influence development teams and vendors without direct authority and to drive reliability improvements through to completion
- Effective use of AI tools in operations work, with the judgment to verify their output
- Pluses: Rancher or similar Kubernetes management platforms; VMware; disaster recovery exercises; knowledge of BDDK regulations; load testing; CKA or CKS certification
Join Us
At ColendiBank, you’ll play a key role in redefining the retail banking experience. We offer opportunities for professional growth in a collaborative, inclusive environment, contributing to a pioneering journey in the digital banking landscape.
Become a core part of our journey. Join ColendiBank as we build the future of digital banking.
ColendiBank is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search