Back to search
PT. Indosat Tbk Linkedin · Posted 10d ago

Head of LLM Evaluation

Jakarta

Linkedin
Continue to application Add your email once, then Caio opens the original posting.

Indexed description

Key Responsibilities

  1. Define Model Evaluation Strategy
  • Design a Comprehensive Evaluation Framework: Develop and maintain a standardized evaluation framework that measures model performance across every stage of the LLM lifecycle, from pre-training to production deployment.
  • Define Quality Metrics and Success Criteria: Establish clear Key Performance Indicators (KPIs) and acceptance thresholds that determine whether a model is ready to progress to the next development stage or production deployment.
  • Develop Indonesia-Centric Evaluation Benchmarks: Create and maintain benchmark datasets that accurately measure the model's understanding of Bahasa Indonesia, regional dialects, cultural context, local regulations, and industry-specific terminology.
  • Define Quality Gates: Define the evaluation process that every model must successfully complete before being approved for deployment.

Benchmark Development

  • Design Comprehensive Benchmark Suites: Develop and maintain standardized benchmark datasets that evaluate the model across a broad range of capabilities, ensuring consistent measurement of performance throughout the model lifecycle.
  • Build Automated Benchmarking Pipelines: Design and implement automated evaluation pipelines that execute benchmark tests whenever a new model, fine-tuned checkpoint, or release candidate is available.
  • Maintain Benchmark Quality and Continuous Improvement: Continuously expand, validate, and improve benchmark datasets to ensure they remain representative of real-world user interactions and emerging AI use cases.

Creating Evaluation Programs

  • Human evaluation program: Design and manage structured human evaluation processes to assess model responses for accuracy, helpfulness, safety, cultural relevance, and overall user experience, ensuring continuous improvement through expert and user feedback.

  • Automated evaluation: Design and implement automated evaluation pipelines that continuously assess model performance across predefined benchmarks, safety tests, regression checks, and quality metrics for every training iteration and production release.

  • Safety evaluation: Design and execute comprehensive safety assessments to evaluate the model's robustness against hallucinations, harmful content, bias, prompt injection, jailbreak attacks, privacy risks, and regulatory compliance before production deployment.

  • Prompt evaluation: Design and execute comprehensive prompt-based test suites to systematically evaluate the model's performance across reasoning, instruction following, multilingual understanding, safety, domain-specific tasks, and real-world user scenarios.

Creating a Continuous Improvement framework

  • Analyze Model Performance Trends: Continuously monitor benchmark results, production metrics, and user feedback to identify performance regressions, weaknesses, and improvement opportunities.


  • Drive Evaluation Enhancements: Regularly refine evaluation datasets, benchmarks, prompts, and testing methodologies to ensure comprehensive coverage of emerging use cases and evolving AI capabilities.


  • Provide Actionable Improvement Recommendations: Collaborate with Research, Data, and MLOps teams by translating evaluation insights into prioritized recommendations that improve model quality, safety, and real-world performance.


Technical Skills Required


Large Language Models (LLMs - Llama, Gemma) : Expert

Natural Language Processing (NLP) :Expert

Deep Learning Frameworks (PyTorch, TensorFlow) : Expert

Data Engineering & Pipeline Development : Expert

Python & Data Science Libraries : Expert

Multilingual & Multimodal AI : Advanced

Model Fine-Tuning & Optimization : Advanced

MLOps & Model Deployment : Advanced

Token-based Monetization Strategies : Intermediate


Qualification & Experience

Education

  • Required: Master’s degree in Computer Science, AI, Data Science, or a related quantitative field.
  • Preferred: PhD with a research focus on NLP, LLMs, or computational linguistics.


Experience

Required:

  • 10+ years of experience in AI/ML, with at least 7 years in a senior leadership role managing AI research and/or development teams.
  • Proven track record of leading the end-to-end development and successful deployment of large-scale machine learning models, specifically LLMs.
  • Extensive hands-on experience with the entire model lifecycle, including pre-training, fine-tuning, RLHF, and evaluation on massive datasets.
  • Strong background in data engineering, including curating, cleaning, and processing large-scale unstructured text data for model training.
  • Experience in defining and implementing monetization strategies for AI services, such as API-based models or “Token as a service.”

Preferred:

  • Specific, demonstrable experience in building or adapting language models for non-English languages, ideally Indonesian or related languages.
  • A history of successful collaboration with platform engineering and product teams to integrate complex AI models into production environments.
  • A strong portfolio of relevant publications in top-tier AI conferences (e.g., NeurIPS, ICML, ACL) or significant contributions to major open-source AI projects.

Free. 20 seconds. No password. See every match in this search.

Create a free Caio profile to unlock more results and save your role and location preferences.

Unlock free search
Want help applying to roles like this? Search Caio for free. If repetitive applications get heavy, Managed Job Search adds supervised execution for $99/month.
View Managed Job Search