Data Scientist
Indexed description
Job Title: Data Scientist
Skills: Artificial intelligence, Machine Learning, NLP, Gen AI, Python, Rest API, Agentic AI, LLM, RAG, Devops and AWS
Experience: 4+ years
Location: Pune and Hyderabad
Duration: Full time
We at Coforge are hiring for Data Scientist role with following skill sets:
LLM & Generative AI
- Design, build, and deploy LLM-powered applications using frameworks such as LangChain, LlamaIndex, or OpenAI API.
- Develop and optimize prompt engineering strategies (few-shot, chain-of-thought, RAG) to improve the accuracy, consistency, and reliability of LLM outputs.
- Implement Retrieval-Augmented Generation (RAG) pipelines using vector databases (e.g., FAISS, Pinecone, Chroma, Weaviate).
- Fine-tune pre-trained LLMs (e.g., GPT, LLaMA, Mistral, Falcon, Claude,Gemini) on domain-specific datasets.
- Validate and structure LLM outputs using Pydantic models and output parsers to ensure data integrity.
Natural Language Processing (NLP)
- Build end-to-end NLP pipelines for real-world tasks including:
- Named Entity Recognition (NER)
- Text Classification & Sentiment Analysis
- Information & Data Extraction from Documents
- Document Summarization & Question Answering
- Semantic Search & Document Similarity
- Work with the Hugging Face Transformers ecosystem to leverage and fine-tune pre-trained models (BERT, RoBERTa, T5, etc.).
- Process large-scale unstructured text data from various sources such as PDFs, emails, scanned documents (OCR), and web content.
Anomaly Detection
- Design and implement anomaly detection systems for various domains, including:
- Financial fraud detection (unusual transactions, payment anomalies).
- Operational anomalies (system logs, network traffic, sensor data).
- Text-based anomalies (unusual document patterns, suspicious NLP signals).
- Apply a wide range of anomaly detection techniques including:
- Statistical Methods: Z-score, IQR, CUSUM.
- ML-based Methods: Isolation Forest, One-Class SVM, Local Outlier Factor (LOF).
- Deep Learning Methods: Autoencoders, LSTM-based sequence anomaly detection, Variational Autoencoders (VAEs).
- Time-Series Methods: ARIMA, Prophet, Seasonal Decomposition.
- Build real-time and batch anomaly detection pipelines that can scale to large datasets.
- Define and tune detection thresholds and alert mechanisms in collaboration with business and operations teams.
Machine Learning (ML)
- Design, train, evaluate, and deploy supervised and unsupervised machine learning models.
- Perform feature engineering, model selection, hyperparameter tuning, and cross-validation.
- Build and maintain end-to-end ML pipelines from data ingestion to model serving.
- Monitor model performance in production and implement retraining strategies to address data drift and model decay.
- Communicate model results, performance metrics, and business impact to technical and non-technical stakeholders.
Python & Software Engineering
- Write clean, modular, production-quality, and well-documented Python code.
- Build and expose ML models as REST APIs using FastAPI or Flask.
- Collaborate with MLOps/DevOps engineers to containerize (Docker) and deploy models in cloud environments.
- Follow best practices in version control (Git), testing, and CI/CD pipelines.
Data & Analytics
- Perform Exploratory Data Analysis (EDA) on structured and unstructured datasets to identify patterns, trends, and anomalies.
- Work with data from relational databases (SQL), data lakes, and cloud storage solutions.
- Create compelling and clear data visualizations (Matplotlib, Seaborn, Plotly) to communicate findings.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search