Senior Data Engineer
Indexed description
Data Engineer
Development, Cloud Engineering & DevOps
Cloud Platform: AWS (Compute, Storage, Networking, IAM and related cloud resources)
Data Engineering Skills: Python, PySpark, SQL, Prefect, Snowflake, Snowpark, dbt, airbyte, Databricks on AWS
DevOps Skills: CI/CD (AWS CodePipeline/CodeBuild/CodeDeploy), Terraform (IaC), Docker, Amazon CloudWatch, Secrets Management, Release & Environment Management
Must-Have Skills
- B.E./B.Tech degree in Computer Science, Engineering, or a related field, with 8-10 years of overall work experience.
- 5+ years of hands-on development experience building Cloud Data Platform Data Engineering solutions covering Data Ingestion, Data Quality Validations, Data Processing and Data Integration.
- Strong hands-on coding experience in Python, PySpark, and SparkSQL, with solid software engineering practices including testing, version control and code reviews.
- Hands-on experience in Airbyte, Snowflake and dbt.
- Hands-on experience with Prefect for workflow orchestration, including designing flows and tasks, scheduling, deployments, and parameterized runs.
- Ability to build reliable, observable Prefect workflows with retry logic, failure handling, and rerun/recovery support for production pipelines.
- Hands-on experience with Amazon S3-based data lakes, Databricks on AWS and other AWS-based data ecosystem services.
- Experience monitoring and troubleshooting orchestrated workflows via the Prefect UI/Cloud, including work pools, deployments and run history.
- Demonstrated willingness and ability to set up and own DevOps practices for the platforms you build (see DevOps Skills below).
- Proficiency in analytics use-case analysis, source system analysis, and data quality assessment.
- Experience coordinating/collaborating with on-shore and off-shore teams for solution delivery.
- Excellent communication and presentation skills.
DevOps Skills
- Design and build CI/CD pipelines using AWS CodePipeline, CodeBuild, and CodeDeploy (or equivalent tools such as GitHub Actions/Jenkins) to automate build, test, and release cycles.
- Provision and manage AWS cloud resources using Infrastructure as Code (Terraform), including version-controlled, reusable modules.
- Containerize applications and workflows with Docker; deploy and manage containers using Amazon ECS/EKS.
- Manage promotion of code and configuration across Development, Test, and Production environments with clear release and rollback strategies (blue-green/canary deployments).
- Implement secure secrets and configuration management using AWS Secrets Manager or Parameter Store, with least-privilege IAM policies.
- Set up monitoring, logging, and alerting using Amazon CloudWatch (metrics, dashboards, alarms) to maintain operational visibility into pipelines and workflows.
- Define and implement retry, recovery, and rerun strategies for failed jobs/workflows to ensure production reliability.
- Write automation scripts (Python/Bash) for deployment, operational tasks, and routine platform maintenance.
- Collaborate with data engineers, application developers, and cloud engineers to support end-to-end delivery, and troubleshoot production issues when required.
Nice-to-Have Skills
- Experience with Snowpark, and additional orchestration frameworks such as Amazon MWAA (Managed Airflow) or AWS Step Functions.
- Exposure to Docker and containerized deployments; familiarity with Amazon ECS/EKS (Kubernetes) is a plus.
- Experience with event-driven architectures (Amazon EventBridge, SQS, SNS, Lambda) and REST API integration across distributed services.
- Exposure to Data Management principles including Data Governance, Data Cataloging, and Master Data Management.
- Exposure to cloud-native MDM tooling.
- Exposure to BI analytics using Amazon QuickSight.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search