Principal Data & AI Platform Architect – Azure Databricks
Indexed description
Responsibilities
Data Pipeline Development & Orchestration
- Design, build, and optimize batch and streaming ETL/ELT pipelines and reusable ingestion frameworks using Azure Data Factory and Databricks across APIs, databases, SaaS platforms, and internal systems.
- Build scalable Delta Lake transformation frameworks using medallion architecture, Spark, and SQL.
- Implement CI/CD, parameterization, triggers, and pipeline automation best practices.
- Architect, manage, and optimize enterprise data environments across Azure Data Lake Storage Gen2 (ADLS Gen2), Azure SQL, and Databricks, including serverless and classic compute strategies, cost governance, and workload isolation strategies.
- Implement DataOps practices including testing, version control, monitoring, and documentation.
- Design and administer the enterprise Unity Catalog structure, including catalogs, schemas, external locations, storage credentials, groups, service principals, and ownership models.
- Implement least-privilege access, governed tags, attribute-based access-control policies, row-level filters, and column-level masking for PHI, PII, financial, and other sensitive information.
- Establish standards for data classification, lineage, auditability, stewardship, retention, certification, and access reviews.
- Partner with Security, Compliance, Privacy, and business data owners to ensure data solutions align with HIPAA and organizational security requirements.
- Govern tables, views, volumes, functions, metric views, dashboards, models, and Genie Agents through Unity Catalog.
- Design and develop Databricks AI/BI Dashboards and domain-specific Genie Agents for clinical, operational, financial, RCM, marketing, and executive use cases.
- Configure trusted datasets, joins, business terminology, instructions, example queries, dimensions, measures, synonyms, and approved KPI definitions.
- Develop and govern Unity Catalog metric views and semantic definitions to ensure consistent reporting across Databricks AI/BI and Power BI.
- Establish testing and monitoring processes for Genie Agent accuracy, data grounding, security, explainability, performance, and user adoption.
- Ensure that AI-generated results respect Unity Catalog permissions and approved business definitions.
- Design and implement enterprise-grade Databricks Lakehouse Medallion architecture (Bronze, Silver, Gold layers).
- Define and enforce data engineering standards, naming conventions, and architectural patterns across all pipelines.
- Lead the architecture of Delta Lake design patterns, including partitioning, optimization, and data lifecycle management.
- Establish scalable serverless and classic compute strategies, job orchestration frameworks, and workspace organization.
- Evaluate and implement new Databricks capabilities and ensure alignment with enterprise data strategy.
- Work closely with clinical, sales, marketing, finance, RCM, operations, HR, and IT teams to understand business needs.
- Provide technical guidance on data engineering patterns and platform capabilities.
- Clearly communicate progress, risks, and technical decisions to data stakeholders and leadership.
- Develop conceptual, logical, dimensional, and physical data models supporting clinical, operational, financial, marketing, RCM, and workforce analytics.
- Establish conformed dimensions, master and reference data standards, governed KPIs, and reusable semantic definitions.
- Reduce conflicting calculations and duplicate business logic across Databricks AI/BI, Power BI, and downstream applications.
- Establish data-quality rules, reconciliation controls, data SLAs, pipeline observability, alerting, incident response, and root-cause-analysis processes.
- Define standards for schema evolution, change data capture, late-arriving data, historical tracking, retries, and recovery.
- Monitor and optimize Databricks consumption using system tables, workload tagging, budget controls, query profiling, serverless and classic compute selection, SQL warehouse configuration, and data-layout optimization.
- Establish cost allocation and accountability by environment, data product, department, and workload.
- 8+ years of progressive experience in data engineering, data architecture, analytics engineering, or cloud data platforms, including demonstrated ownership of production Databricks architecture.
- Demonstrated technical leadership and mentoring experience within data engineering or architecture teams.
- Hands-on expertise with Azure Data Factory, including pipelines, mapping data flows, Integration Runtime configuration and management, triggers, and monitoring.
- Hands-on expertise with Azure Databricks, including notebooks, Apache Spark, Delta Lake, Databricks SQL, Lakeflow Jobs, Lakeflow pipelines, and workflow orchestration.
- Advanced SQL expertise, including complex transformations, dimensional and semantic modeling, query-plan analysis, Delta Lake optimization, Databricks SQL, and SQL warehouse performance tuning.
- Advanced proficiency in Python and PySpark for data engineering, reusable framework development, automation, testing, and performance optimization.
- Deep expertise in Medallion Lakehouse architecture (Bronze/Silver/Gold) and Delta Lake optimization techniques.
- Demonstrated experience designing, implementing, and operating enterprise Databricks environments across development, testing, and production, including security, governance, deployment, performance, and cost management responsibilities.
- Strong understanding of Databricks Unity Catalog, data governance, and security models.
- Strong understanding of HIPAA, PHI/PII safeguards, least-privilege access, data retention, auditability, and secure healthcare data integration.
- Experience defining data platform standards, frameworks, and best practices
- Experience with AI/ML workflows, feature engineering, or model enablement.
- Experience integrating data across EHR/EMR, CRM, patient-engagement, contact-center, marketing, finance/ERP, HRIS, and revenue-cycle platforms.
- Experience designing enterprise data models and governed KPIs for healthcare operations, including patient volume, referrals, scheduling, conversion, provider productivity, revenue cycle, payer performance, denials, collections, labor, and clinic-level financial performance.
- Familiarity with real-time processing (Structured Streaming) within Databricks.
- Experience with master data management, reference data, and entity-resolution strategies across patients, providers, locations, payers, legal entities, and acquired practices.
United Vein & Vascular Centers (UVVC) is distinguished by its innovative approach to diagnosing and treating a variety of vascular conditions that affect the pelvis and lower extremities. With a team of committed specialists, cutting-edge medical technology, and a patient-centric approach that emphasizes minimally invasive procedures, UVVC ensures superior care and optimal outcomes for it’s patients.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search