Back to search
Aarista Technologies Linkedin · Posted 10d ago

Senior Databricks Engineer - Revenue Cycle Management (RCM)

Houston, Texas, United States

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

About the Role


We are seeking a Senior Databricks Engineer to design, build, and optimize large-scale data pipelines powering our healthcare Revenue Cycle Management platform. You will own end-to-end data engineering across EDI claim processing (837, 835, 277, 999), remittance workflows, AI-driven medical coding, and real-time operational data delivery using the Medallion architecture on Azure Databricks.


Key Responsibilities:


  • Design & maintain production data pipelines using PySpark and Scala on Azure Databricks, processing healthcare EDI transactions (837 Claims, 835 Remittances, 277/999 Acknowledgments)
  • Implement and optimize Delta Lake tables with ACID transactions, Z-ordering, partitioning strategies, and data versioning (time travel) for petabyte-scale healthcare data
  • Build MongoDB sync pipelines to deliver Gold-layer data into operational MongoDB collections for real-time application access
  • Develop/Maintain Scala-based reporting and workflow orchestration engines, including dynamic report generation, AP analysis, and notification triggers
  • Integrate with Azure ecosystem: Azure Blob Storage, Azure Data Lake, Azure AD, Azure DevOps CI/CD pipelines, and Databricks Unity Catalog
  • Manage multi-tenant data architectures with per-client configurations, secrets management (Databricks secrets scopes), and feature flag-driven rollouts
  • Build and maintain SQL Server integrations via JDBC for workflow metadata, feature flags, report configurations, and state management
  • Implement data quality frameworks with validation layers at each pipeline stage (patient info, diagnosis codes, CPT codes, facility mappings, rendering providers)
  • Write ad-hoc data correction, migration, and backfill scripts for production data reconciliation
  • Participate in code reviews, PR validation pipelines, and engineering best practices (commit/lint, conventional commits)


Required Skills & Qualifications:


  • Core: 6–7 years of data engineering experience with 3+ years on Databricks/Spark
  • Languages: Proficient in PySpark (primary) and Scala/Spark; Python 3.11+
  • Data Lake: Deep expertise in Delta Lake — medallion architecture, ACID transactions, schema evolution, Z-ordering, partition pruning
  • Cloud: Strong Azure experience: Blob Storage, Data Lake Gen2, Azure AD, Key Vault, Azure DevOps
  • Databases: Production experience with MongoDB (aggregation pipelines, indexing, bulk sync) and SQL Server (JDBC, stored procedures)
  • EDI/Healthcare: Familiarity with HIPAA EDI standards (X12 837, 835, 277, 999) or willingness to learn quickly
  • Orchestration: Databricks Workflows, notebook orchestration, parameterized jobs
  • DevOps: Azure Pipelines (YAML), CI/CD for Databricks artifacts, PR validation
  • Data Quality: Experience implementing validation frameworks, data reconciliation, and error tracking


Preferred Qualifications:


  • Experience in healthcare/RCM domain — claims lifecycle, payer/provider workflows, CPT/ICD-10 coding would be a plus
  • Hands-on with Databricks Unity Catalog for data governance
  • Exposure to AI/ML integration in data pipelines — LLM orchestration, model observability (Langfuse or similar)
  • Experience with feature flag systems for gradual data pipeline rollouts
  • Multi-tenant SaaS data architecture design
  • Performance tuning Spark jobs at scale (shuffle optimization, broadcast joins, adaptive query execution)
  • Tech Stack Summary
  • Azure Databricks · PySpark · Scala/Spark · Delta Lake · MongoDB · SQL Server · Azure Blob Storage · Azure DevOps · Unity Catalog · EDI X12 (837/835/277/999) · Langfuse · Python 3.11 · JDBC · Azure AD


What You'll Impact:


  • Process thousands of healthcare claims daily across multiple clients
  • Reduce claim processing time through pipeline optimization
  • Enable AI-powered medical coding that improves coder productivity
  • Deliver real-time operational dashboards for revenue cycle teams
  • Ensure HIPAA-compliant data processing at every layer
Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search