Databricks Developer
Indexed description
Why Kumaran?
We blend engineering discipline with AI innovation, helping clients modernise with confidence, automate with clarity, and scale with purpose. Our global delivery model ensures agility, responsiveness, and seamless collaboration, with clients always at the heart of every engagement. At Kumaran, we don’t just solve problems, we engineer future-ready transformations.
Job Summary
We are seeking a skilled Databricks Developer to design, develop, and optimize scalable data pipelines and analytics solutions using Apache Spark, Python (PySpark), and SQL within modern cloud environments. The ideal candidate will have hands-on experience working with Databricks, Delta Lake, and cloud data platforms, and will be responsible for building high-performance data processing workflows that support data-driven decision-making.
Key Responsibilities
Pipeline Development
- Design, develop, and maintain ETL/ELT data pipelines using PySpark and SQL in Databricks notebooks.
- Process large-scale datasets and ensure reliable and efficient data transformations.
- Implement data lake and data warehouse architectures using Databricks Delta Lake and Delta Live Tables.
- Build and manage Medallion Architecture (Bronze, Silver, Gold layers) for structured data processing.
- Optimize Spark jobs and queries for performance, scalability, and cost-efficiency.
- Manage cluster configurations, partitioning strategies, and caching mechanisms.
- Integrate Databricks solutions with cloud services such as:
- Azure Data Factory (ADF)
- Azure Data Lake Storage (ADLS) Gen2
- AWS S3 or other cloud storage platforms
- Implement data quality checks, validation frameworks, and monitoring processes.
- Ensure data security, encryption, masking, and lineage tracking.
- Build and manage automated workflows and scheduling using Databricks Jobs, Airflow, or CI/CD pipelines.
- Integrate with DevOps tools such as Azure DevOps or Jenkins for continuous integration and deployment.
- Work closely with data engineers, analysts, and business stakeholders to translate business requirements into scalable data solutions.
- Participate in Agile development processes including sprint planning and technical discussions.
- 3–5+ years of experience with Databricks, Apache Spark, and Python (PySpark).
- Strong experience building scalable ETL/ELT pipelines.
- Advanced knowledge of Spark SQL or Databricks SQL for data transformation and analysis.
- Hands-on experience with Azure Databricks, AWS, or Google Cloud Platform (GCP).
- Experience with Delta Lake and Medallion Architecture (Bronze/Silver/Gold layers).
- Strong understanding of data modeling and data warehouse concepts.
- Proficiency with Git and CI/CD tools such as Azure DevOps or Jenkins.
- Experience with Data Build Tool (dbt).
- Knowledge of workflow orchestration tools such as Apache Airflow.
- Experience with Machine Learning workflows using MLflow.
- Familiarity with data governance and enterprise data platforms.
- Bachelor’s or Master’s degree in Computer Science, Information Technology, Data Engineering, or a related field.
- Strong problem-solving and analytical skills
- Ability to handle large-scale data processing environments
- Good communication and collaboration skills in Agile teams
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search