Back to search
Thompsons HR Consulting Pvt Ltd Linkedin · Posted 27d ago

Data Scientist/ NLP

India

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Data Scientist NLP Job Overview

We are looking for a hands-on Data Scientist specializing in Natural Language Processing (NLP) to design, develop, evaluate, and deploy production-grade NLP and machine learning solutions for complex, text-driven workflows.

The ideal candidate should have strong expertise in Python, SQL, NLP, Transformers, embeddings, semantic search, information retrieval, and machine learning, with the ability to take solutions from experimentation through production deployment.

Key Responsibilities

  • Design and develop NLP and machine learning pipelines to process noisy, heterogeneous text data and transform it into clean semantic representations for modeling, retrieval, analytics, and downstream applications.
  • Build and optimize semantic search and retrieval systems using embeddings, vector databases, similarity search, and ranking techniques.
  • Develop solutions for candidate ranking, out-of-vocabulary handling, semantic matching, and information discovery.
  • Design, implement, and evaluate supervised and hybrid ML approaches, including:
    • Multi-output classification
    • Hierarchical classification
    • Named Entity Recognition (NER)
    • Entity extraction and parsing
    • Clustering
    • Rule-based + ML hybrid systems
  • Work with transformer-based models for text understanding, classification, similarity, extraction, and retrieval use cases.
  • Perform detailed analysis of relationships and decision boundaries across free-text fields using:
    • Conditional distributions
    • Entropy
    • Mutual information
    • Directional association
    • Embeddings
    • Predictive ablation studies
  • Design experiments and establish appropriate model evaluation metrics and benchmarks.
  • Compare different models and approaches based on accuracy, performance, scalability, latency, and business impact.
  • Fine-tune and evaluate transformer models using frameworks such as Hugging Face and SentenceTransformers.
  • Deploy, monitor, troubleshoot, and continuously improve production ML/NLP services.
  • Collaborate closely with platform, backend, data engineering, and product teams to integrate ML solutions into production systems.
  • Communicate technical findings, model performance, and trade-offs effectively to both technical and non-technical stakeholders.
Required Skills

  • Strong programming experience in Python and SQL.
  • Hands-on experience developing production-grade data pipelines and machine learning workflows.
  • Strong understanding of Natural Language Processing (NLP) and text analytics.
  • Practical experience with one or more of the following:
    • Text Classification
    • Semantic Similarity
    • Text Embeddings
    • Information Retrieval
    • Search & Ranking
    • Clustering
    • Named Entity Recognition (NER)
    • Entity Extraction
  • Hands-on experience with:
    • Hugging Face
    • SentenceTransformers
    • Tokenization
    • Transformer-based models
    • Model fine-tuning
    • Model evaluation
  • Strong understanding of embeddings, vector search, and similarity search.
  • Knowledge of cosine similarity and Approximate Nearest Neighbor (ANN) search methods.
  • Familiarity with information retrieval metrics such as:
    • Recall@K
    • MRR (Mean Reciprocal Rank)
    • NDCG (Normalized Discounted Cumulative Gain)
  • Strong analytical and problem-solving skills with the ability to design experiments, define evaluation metrics, and interpret model results.
  • Ability to evaluate model trade-offs and clearly communicate technical findings.
Nice to Have

  • Experience in healthcare, medical imaging, document intelligence, enterprise search, recommendation systems, knowledge retrieval, or routing systems.
  • Knowledge of healthcare and enterprise data standards such as:
    • DICOM
    • PACS/RIS
    • HL7
    • FHIR
  • Experience with MLOps and production ML systems.
  • Experience with cloud platforms and API-based model deployment.
  • Experience with model serving, monitoring, logging, and performance optimization.
  • Knowledge of Responsible AI, data privacy, and secure handling of sensitive text data.
  • Experience working with vector databases/search platforms and large-scale retrieval systems.
Preferred Candidate Profile

The ideal candidate will have a combination of NLP expertise, machine learning fundamentals, information retrieval knowledge, and production engineering experience. Candidates with experience building semantic search, embedding-based retrieval, classification, RAG, document intelligence, or enterprise search solutions will be highly preferred.

Core Skills

Python | SQL | NLP | Machine Learning | Hugging Face | SentenceTransformers | Transformers | Embeddings | Semantic Search | Vector Search | Information Retrieval | Ranking | Text Classification | NER | Clustering | Model Fine-tuning | Recall@K | MRR | NDCG

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search