Python Engineer — Evaluator Library
Indexed description
Intro:
We are looking for a Python Engineer — Evaluator Library to design and implement reusable evaluation components that ensure the quality, safety, and compliance of enterprise AI agents and LLM-powered workflows. You will build custom evaluation capabilities used across AI platforms, focusing on automated quality validation, workflow compliance, PII protection, and structured output verification. Working closely with AI Platform Engineers, ML Engineers, and DevOps teams, you will help establish reliable evaluation standards and scalable quality assurance mechanisms for agent-based systems.
Project Overview:
Our customer is a multinational corporation with more than a century of history and offices in over 180 countries. Their most ambitious goal at the time is to introduce a range of Reduced-Risk Products (RRPs). The target audience is more than 1 billion consumers around the globe. IT platform hosts 700+ applications.
Intellia's mission is to help the client with the engineering of a comprehensive software ecosystem for a game-changing IoT product on the margin of innovative consumer experience and cutting-edge technology. Our teams are involved in the engineering of core platform components for best-in-class eCommerce, Digital Marketing and IoT solutions. As an Engineer, you will become a part of Core Architecture Team and be responsible for the architecture, implementation of best practices in our Digital Engineering Enterprise Platform.
The Platform is a set of services and internet applications that accelerate the development and delivery of software applications by taking care of common SDLC challenges. The Platform provides access and consumption for engineering teams to a set of services, technologies, practices for their development and for operating their application, ensuring a set of compliance and best practices.
Requirements:
Skills:
• Python (Lambda functions as AWS AgentCore custom code-based evaluators)
• LLM-as-judge prompt engineering for subjective evaluation dimensions
• PII detection (regex-based + AWS Bedrock Guardrails)
• Tool response schema validation (JSON Schema — TOOL_CALL level evaluator)
• Workflow contract compliance checking (SESSION level evaluator)
• Numerical accuracy validation logic (TRACE level evaluator)
Experience:
• 4+ years Python engineering
• LLM evaluation or quality assurance for AI/ML systems
• AWS Lambda function development and deployment
Nice-to-have
• AWS AgentCore Evaluation custom evaluator Lambda registration
• AWS Bedrock Guardrails for PII detection integration
• CloudWatch Logs as evaluator output sink
Responsibilities:
- Design, develop, and maintain reusable Python-based evaluator libraries for AI agents and LLM-powered workflows.
- Implement AWS Lambda-based custom evaluators to perform deterministic quality, compliance, and validation checks.
- Develop LLM-as-a-judge evaluation logic to assess subjective dimensions such as relevance, helpfulness, consistency, and response quality.
- Build automated PII detection evaluators using regex-based techniques and AWS Bedrock Guardrails integrations.
- Implement TOOL_CALL-level validation mechanisms to verify structured outputs, JSON schema compliance, and tool response correctness.
- Develop SESSION-level evaluators to validate workflow contract compliance, execution integrity, and cross-step behavioral expectations.
- Create TRACE-level evaluators for numerical accuracy verification, calculation consistency, and deterministic result validation.
- Integrate evaluator components with AWS AgentCore Evaluation workflows and enterprise AI quality pipelines.
- Collaborate with AI platform teams to define evaluation standards, scoring methodologies, and quality acceptance criteria.
- Design and maintain unit, integration, and validation tests for evaluator libraries and quality frameworks.
- Support observability and troubleshooting by integrating evaluation outputs with CloudWatch logging and monitoring capabilities.
- Contribute to enterprise AI governance initiatives by improving evaluation coverage, auditability, and compliance controls.
Create a free Caio profile to unlock more results and save your role and location preferences.
Unlock free search