Data Scientist
Indexed description
Location: Fairfield County, CT| Work Setup: Hybrid | Compensation: $180,000 – $200,000 Base + Bonus & Benefits
The Opportunity
We are seeking a versatile Data Scientist who thrives at the intersection of complex data engineering and advanced quantitative research. In this role, you will act as the vital connector between technical research teams and data infrastructure, building the reliable pipelines, features, and datasets that power sophisticated machine learning models.
If you are obsessed with data integrity, love tackling unstructured datasets, and want your work directly tied to core production strategies, this position is built for you.
Core Responsibilities
- Collaborative Data Modeling: Partner closely with research partners to understand project specs, building tailored, model-ready features and datasets for active machine learning initiatives.
- Complex Data Transformation: Ingest, clean, and map disparate structured and unstructured data sources. Tackle hard entity-matching, tagging, and linkage problems across historical signals and textual assets.
- Automated Quality & Profiling: Design end-to-end data validation frameworks using automated scripts, LLM-driven verification steps, and targeted manual audits to guarantee 100% data traceability.
- Entity Resolution & Mapping: Construct point-in-time knowledge graphs and mapping schemas to handle complex historical events like corporate restructuring, M&A, and public offerings.
- Alternative Data Exploration: Onboard, evaluate, and sanitize new alternative datasets to unlock novel insights for active research streams.
Qualifications & Stack
- 4+ Years Experience: Proven track record in data science, quantitative data engineering, or a heavily research-driven environment.
- Technical Core: Advanced Python (pandas, NumPy) and production SQL (PostgreSQL).
- Engineering Rigor: Hands-on experience with version control (Git), testing frameworks (PyTest), API development, and CI/CD pipelines.
- Data Mastery: Demonstrated experience handling both clean tabular data and messy, unstructured text with a sharp eye for edge cases and precise documentation.
Bonus Points For
- Practical experience applying LLM APIs (e.g., Claude Code, Bedrock, Codex) for automated validation, featurization, or entity resolution.
- Background in financial datasets, systematic strategy research, or market data.
- Familiarity with scikit-learn, PyTorch, AWS infrastructure (S3, Batch), or distributed computing tools.
What We Look For In You
- Meticulous Mindset: You treat data accuracy as non-negotiable and take pride in rock-solid pipeline reliability.
- Clear Communicator: You easily bridge the gap between technical developers and domain-expert researchers.
- Adaptable Execution: You navigate multi-project environments comfortably as team goals and technologies evolve.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search