Senior DevOps Engineer / Site Reliability Engineer (SRE)
Indexed description
Job Overview:-
We are seeking an experienced Senior DevOps Engineer / Site Reliability Engineer (SRE) to design, build, and operate a scalable cloud-based data platform. You will be responsible for managing enterprise-scale Databricks environments, automating infrastructure, improving platform reliability, and enabling data teams through self-service solutions.
Key Responsibilities
- Manage and administer Databricks environments across multiple workspaces.
- Design and implement scalable, secure multi-tenant platform architectures.
- Automate infrastructure provisioning using Terraform and Infrastructure as Code (IaC).
- Build and maintain CI/CD pipelines for platform deployments.
- Develop automation tools using Python and Bash/Shell scripting.
- Configure and manage cloud infrastructure on Google Cloud Platform (GCP).
- Implement monitoring, logging, alerting, and incident response processes.
- Define and maintain SLOs, SLAs, and platform reliability standards.
- Collaborate with Data Engineering, Security, and Product teams.
- Improve cloud cost optimization and platform governance.
- Utilize AI-assisted development tools such as GitHub Copilot and ChatGPT to enhance engineering productivity.
Required Skills
- 5+ years of experience in DevOps, Platform Engineering, Infrastructure Engineering, or Data Platform Engineering.
- Hands-on experience with Databricks administration, Unity Catalog, Cluster Policies, and Workspace Management.
- Strong experience with Google Cloud Platform (GCP), including IAM, VPC, GCS, and BigQuery.
- Expertise in Terraform and Infrastructure as Code (IaC).
- Strong scripting skills in Python and Bash/Shell.
- Experience with CI/CD tools such as GitHub Actions, GitLab CI, or Cloud Build.
- Knowledge of cloud security, RBAC, secrets management, and audit logging.
- Strong troubleshooting, communication, and stakeholder management skills.
- Experience working with DevOps and Site Reliability Engineering (SRE) practices.
Preferred Skills
- Apache Spark
- Databricks Workflows
- Databricks Asset Bundles
- Apache Airflow
- Grafana
- Datadog
- Cloud Monitoring
- FinOps
- MLOps
- Delta Live Tables
- Data Platform Governance
Requirements
- Business-level English communication skills.
- Experience working in cloud-native and enterprise-scale environments.
- Candidates currently residing in Japan with a valid work visa are preferred.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search