Back to search
Real Linkedin · Posted yesterday

Founding Data Science Engineer (Depth)

Syracuse-Auburn

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

About the Opportunity

Our client is an innovative, early-stage artificial intelligence company focused on transforming how life sciences data is collected, structured, and utilized for advanced reasoning and decision-making. They are seeking a Founding Data Science Engineer to play a critical role in expanding and enriching a proprietary evidence platform that powers next-generation AI applications across biotechnology, pharmaceuticals, and healthcare.

Position Overview

This founding-level hire will be responsible for designing and implementing AI-driven systems that transform complex scientific, clinical, regulatory, and business information into structured, trustworthy, and source-linked data. The ideal candidate will combine expertise in data science, machine learning, data engineering, and AI-driven information extraction with a strong ability to model complex real-world relationships and develop scalable data architectures.

The successful candidate will help define how diverse evidence types are represented, ingested, validated, and leveraged by downstream AI reasoning systems. This position offers significant ownership and influence over product architecture, data strategy, and AI workflow design.

Key Responsibilities

  • Design and develop AI-powered ingestion and extraction systems for scientific publications, clinical trial records, patents, regulatory documents, financial filings, conference materials, press releases, and other public data sources.
  • Create structured data models that capture entities, relationships, provenance, confidence levels, and validation requirements across multiple evidence types.
  • Build and optimize agentic AI workflows that transform unstructured content into traceable, high-quality structured datasets.
  • Define and maintain data quality standards, validation frameworks, benchmarking methodologies, and extraction performance metrics.
  • Develop scalable approaches for entity resolution, data harmonization, and canonicalization across diverse source materials.
  • Partner closely with engineering and domain experts to establish data contracts and ensure downstream AI systems can confidently consume structured evidence.
  • Implement evaluation frameworks including gold-standard data sets, regression testing, precision and recall analysis, and extraction quality monitoring.
Required Qualifications

* Strong experience with Python and modern data engineering practices.

* Experience building machine learning, artificial intelligence, NLP, LLM, or Generative AI solutions in production environments.

* Deep understanding of data modeling, schema design, structured data systems, and knowledge representation.

* Strong analytical thinking and ability to develop practical solutions to ambiguous data challenges.

* Experience designing quality controls, validation methodologies, and performance measurement frameworks.

* Excellent communication skills and ability to translate domain expertise into reproducible systems and workflows.

Preferred Qualifications

* Experience with Large Language Models (LLMs), Agentic AI systems, Retrieval-Augmented Generation (RAG), prompt engineering, or AI evaluation frameworks.

* Background in life sciences, biotechnology, pharmaceuticals, clinical research, healthcare data, or scientific informatics.

* Experience with knowledge graphs, graph databases, semantic data models, ontology development, or entity resolution.

* Familiarity with scientific literature mining, clinical trial data, regulatory information, or research data platforms.

* Experience developing internal review tools, audit workflows, benchmark data sets, or data governance solutions.

Ideal Background

Candidates may come from Data Science, Machine Learning Engineering, Applied AI Engineering, Knowledge Graph Engineering, Data Platform Engineering, Bioinformatics, Computational Biology, Scientific Data Engineering, or related technical disciplines.

Desired Skills and Experience

5+ years of experience in Data Science, Machine Learning, Data Engineering, LLMs/Generative AI, Knowledge Graphs, and Life Sciences data environments.

EOE Statement: Specialist Staffing Group is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or veteran status.


In addition to base pay, direct-hire employees may be eligible for client offered benefits such as medical, dental, and vision coverage, and paid leave where required by applicable law. Eligibility may vary based on factors such as location and hire date and is subject to change.


To find out more about Real, please visit www.realstaffing.com

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search