Back to search
Staffingine LLC Linkedin · Posted yesterday

AI Data Engineer

Princeton, New Jersey, United States

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Staffingine LLC is excited to announce an excellent Fulltime opportunity with one of our top clients!


Position: AI Data Engineer

Location: Princeton, NJ & NYC, NY – Hybrid

Salary: $160,000–$170,000 USD

Experience: 10+ Years

Employment Type: Full-Time


Role Overview

  • The AI Data Engineer will specialize in building and optimizing machine learning data pipelines, focusing on AI model tracking, lifecycle management, and integration with AI governance systems. This role combines data engineering expertise with AI/ML knowledge to support the organization's broader data and AI infrastructure initiatives.


Key Responsibilities

  • Design AI data pipelines for model metadata, data lineage, and performance tracking.
  • Build scalable data infrastructure using Databricks and Spark.
  • Develop MCP servers and AI data distribution solutions.
  • Build feature engineering and data preprocessing pipelines.
  • Implement MLflow for model versioning, experiment tracking, and model registry.
  • Develop workflows for AI agent discovery and inventory management.
  • Build knowledge graphs for AI model relationships and data lineage.
  • Develop real-time pipelines for model monitoring, drift detection, and performance tracking.
  • Establish data quality, security, and compliance frameworks for AI data and model artifacts.
  • Collaborate with Data Scientists, ML Engineers, and Architects on data and feature-store solutions.


Required Skills and Qualifications

  • Bachelor's or Master's degree in Computer Science, Data Science, Machine Learning, or related field.
  • 5–7 years of hands-on experience in data engineering, with at least 2 years focused on AI/ML workloads.
  • Expert proficiency in Python and experience with ML frameworks such as PyTorch, TensorFlow, or scikit-learn.
  • Strong experience with Databricks, Apache Spark, and distributed computing for ML workflows.
  • Deep understanding of the machine learning lifecycle, including model training, deployment, and monitoring processes.
  • Experience with feature engineering, data preprocessing techniques, and ML data pipelines.
  • Knowledge of vector databases, embeddings, and similarity search for AI applications.
  • Proficiency in SQL for structured and unstructured data management.
  • Understanding of data governance, model governance, and AI ethics principles.
  • Strong analytical and problem-solving capabilities with attention to data quality.
  • Excellent collaboration skills for working with data scientists, ML engineers, and architects.


Preferred / Nice-to-Have Skills

  • Experience with generative AI applications, including RAG (Retrieval-Augmented Generation) and fine-tuning.
  • Knowledge of LangChain, HuggingFace, or other GenAI frameworks.
  • Familiarity with Azure ML, AWS SageMaker, or Google Vertex AI platforms.
  • Experience with graph databases (Neo4j, Amazon Neptune) for knowledge graph implementation.
  • Understanding of AI model explainability and interpretability techniques.
  • Experience with A/B testing frameworks for ML model evaluation.
  • Certification in Databricks, AWS, Azure, or GCP AI/ML services.
  • Publications or contributions to open-source ML projects


Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search