Vi
Linkedin · Posted 1mo ago
Senior Data Engineer
Continue to application
Add your email once, then Caio opens the original posting.
Indexed description
- Vi is an enterprise AI platform for health enterprises - healthcare, biopharma, and wellness. We deploy agentic AI and predictive models into production environments where the output drives next best actions for patients, care teams, and operations to deliver ROI and improve health outcomes.
- We are looking for a Senior Data Engineer to build and scale the data foundation behind Vi's platform and products. You will own complex data end-to-end - from ingestion and transformation through modeling, quality, observability, and production delivery.
- This is a hands-on senior IC role for a strong builder who can solve difficult data problems independently, set a high technical bar, and collaborate closely with engineering, DS, and product. You will turn large, fragmented datasets into reliable, reusable capabilities that power every Vi product.
- Build and own scalable data pipelines- Design, implement, and operate robust pipelines for high-volume structured and unstructured data, with validation, monitoring, lineage, and recovery built in.
- Scale the platform for growth- A key near-term initiative is re-architecting the system to support a significantly larger customer base. You will own performance and cost-efficiency across pipelines and services, keeping reliability and operating costs under control as the platform scales.
- Build across the stack- This is not a pipelines-only role. You will also write backend services and some frontend, including the internal backoffice the team runs on. We hire builders, not narrow specialists.
- Own the core data tables- Own schema design and evolution, data contracts, and the modeling standards the team follows — naming, shared dimensions, normalization, documentation. Be accountable when a table is wrong, late, or drifting.
- Level up the team's data work- Pair with and advise software engineers and data scientists on Spark, SQL, and modeling, and help turn notebook-grade code into production-grade pipelines.
- Partner cross-functionally- Translate product, client, compliance, and business requirements into clear technical designs and dependable production systems.
- Spark at scale- You have tuned real Spark jobs for performance and cost — skew, shuffle, partitioning, memory, spill — run pipelines over TB-scale or billions of rows in production, and can reason about the physical execution plan, not just write DataFrame code.
- 5+ years of professional experience building and owning production systems.
- Strong Python and SQL, with maintainable, tested production code.
- Strong software engineering fundamentals across the stack. You can own backend services and pick up frontend when the work needs it — not a pipelines-only specialist.
- AI-first way of working- You build with AI in your day-to-day development, using it to move faster and raise the quality of what you ship.
- Deep experience designing and operating ETL/ELT pipelines, data models, and distributed data-processing systems.
- Comfortable advising and pairing with other engineers and data scientists on data work.
- Strong AWS experience: S3, Glue, EMR, Athena, and related compute and orchestration services.
- Experience with modern data lakehouse or warehouse architectures. Apache Iceberg is a strong advantage.
- Experience with workflow orchestration (Airflow or similar), CI/CD, Docker, Git, and infrastructure as code such as AWS CDK and CloudFormation.
- Strong understanding of data quality, schema evolution, lineage, observability, privacy, security, and access controls. Experience with regulated or sensitive data, such as healthcare / PHI, is an advantage.
- High comfort in a fast-moving environment with incomplete requirements, high ownership, and a strong sense of urgency.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search
Want help applying to roles like this?
Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search