Data Engineer
Indexed description
Data Engineer for Clinical Operations
One of my premier clients in the pharmaceutical industry is seeking a talented Data Engineer to join their Global Drug Development IT division. This role is based in Princeton, New Jersey, and focuses on creating high-performance data environments for cross-study operations and specimen management.
In this position, you will serve as a hands-on technical expert responsible for the entire data product lifecycle. You will work to build secure, scalable, and innovative data ecosystems that support clinical R and D programs, utilizing modern cloud platforms and Generative AI to solve complex data challenges.
Key Responsibilities
Collaborate with clinical study teams, trial managers, and IT partners to drive the adoption of the corporate data platform.
Build and maintain production-grade data pipelines specifically for specimen tracking, biobanking workflows, and cross-trial data aggregation.
Enhance data solutions to improve interoperability and scalability across clinical research programs.
Optimize platform performance and cost-effectiveness using cloud-native processing and Databricks Delta Lake.
Partner with data product owners to handle data modeling, lineage, and access governance.
Implement Databricks Unity Catalog to manage metadata and enforce strict data governance.
Develop and deploy Generative AI and NLP-driven applications to improve specimen traceability and automate compliance.
Utilize RAG, fine-tuning, and vector embeddings to create advanced data discovery tools.
Leverage Mosaic AI and MLflow to manage machine learning models at scale.
Act as a technical mentor for junior staff and external vendors on best practices and business alignment.
Required Qualifications and Experience
At least 2 years of professional experience in Data Engineering, Analytics, or AI and ML within a cloud environment.
Strong technical expertise with Databricks, including Workflows, Unity Catalog, and Delta Lake. Databricks certification is highly preferred.
Deep proficiency in Python, SQL, Spark, and PySpark.
Proven experience with Generative AI frameworks, including LLM architectures, prompt engineering, and RAG.
Background in creating ETL or ELT pipelines and semantic data models for large, complex datasets.
Previous experience in life sciences, clinical trial operations, or biobanking is a significant advantage.
Excellent communication skills with the ability to explain technical concepts to non-technical business stakeholders.
Compensation and Benefits
The estimated starting salary for this role is between 87,810 and 106,399 USD annually, depending on experience and skills.
The package includes eligibility for incentive cash and stock opportunities.
Comprehensive health coverage including medical, dental, and vision.
Financial protections such as a 401k plan, disability insurance, and life insurance.
Wellbeing programs and flexible paid time off to support work-life balance.
If you are a data professional looking to apply your skills toward life-changing pharmaceutical research, I encourage you to apply even if you do not meet every single qualification listed.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search