Senior Database Engineer
Indexed description
We are seeking a highly skilled CockroachDB Database Engineer with strong Site Reliability Engineering (SRE) experience to design, implement, manage, and optimize large-scale distributed database platforms. The ideal candidate will have hands-on expertise in CockroachDB administration, performance tuning, high availability, disaster recovery, automation, observability, and operational reliability. The role requires close collaboration with development, infrastructure, and platform engineering teams to ensure highly available, resilient, and scalable database services.
Key Responsibilities
Design, deploy, administer, and maintain production-grade CockroachDB clusters across cloud and on-premises environments.
Monitor database health, performance, latency, throughput, and resource utilization to ensure service reliability and availability.
Implement and manage backup, restore, disaster recovery, and business continuity strategies.
Perform database capacity planning, performance tuning, indexing, and query optimization.
Develop automation scripts and Infrastructure-as-Code (IaC) solutions to streamline provisioning, upgrades, and operational tasks.
Establish and manage SRE practices including Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets.
Drive incident management, root cause analysis (RCA), postmortems, and preventive remediation activities.
Build and maintain monitoring, logging, and alerting solutions using tools such as Prometheus, Grafana, ELK, Datadog, or similar platforms.
Collaborate with DevOps and Engineering teams to improve platform reliability, scalability, security, and operational excellence.
Support production releases, database migrations, version upgrades, and platform modernization initiatives.
Participate in on-call rotation and provide support for critical production incidents.
Implement database security controls, access governance, auditing, and compliance best practices.
Required Skills
Database Technologies
Strong hands-on experience with CockroachDB Administration
Expertise in distributed SQL databases and cluster management
Database performance tuning and query optimization
Backup, recovery, replication, and data protection strategies
High Availability and Disaster Recovery architecture
Site Reliability Engineering (SRE)
Strong understanding of SRE principles and operational excellence
Experience defining and tracking SLIs, SLOs, and Error Budgets
Incident response, RCA, and reliability engineering practices
Production monitoring, observability, and capacity management
Reliability automation and operational process improvement
Cloud & Automation
Experience with AWS, Azure, or GCP
Infrastructure as Code (Terraform, Ansible, etc.)
Linux/Unix administration
Scripting using Python, Shell, or Go
CI/CD pipeline integration and automation
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search