Join the Enterprise Agent Development Platform project at EPAM. We are building a cloud-native platform that enables engineering teams to develop, deploy, and operate AI agents in production.
The platform combines modern agent frameworks, AWS infrastructure, CI/CD, observability, and governance to make AI development faster and more reliable.
As a Python AI Evaluation Engineer, you will own a key part of this platform: building the capabilities that measure and validate the quality of AI solutions. You will design evaluation approaches, develop custom evaluators, and integrate quality checks into the software delivery lifecycle.
This role is a strong fit for an engineer who enjoys solving new GenAI challenges and turning them into practical, automated engineering solutions.
Responsibilities
- Design and implement evaluation frameworks for LLM-based applications and AI agents
- Develop LLM-as-a-Judge and deterministic, code-based evaluators
- Build custom Python evaluators for quality and behavioral checks
- Define evaluation criteria, metrics, thresholds, and acceptance rules
- Evaluate agent behavior across individual responses, tool calls, and complete workflows
- Work with OpenTelemetry traces and spans as evaluation data
- Integrate evaluations into CI/CD pipelines and automated deployment gates
- Enable continuous quality monitoring of solutions in production
- Establish reusable evaluation patterns and engineering standards
- Work closely with AI Engineers, Architects, and Platform Engineers to embed quality into the development process
Requirements
- 5+ years of experience in ML Engineering, AI Engineering, or AI Platform Engineering
- Strong Python development experience
- Hands-on experience with LLM/GenAI evaluation
- Experience designing and implementing evaluation frameworks
- Experience developing custom or deterministic evaluators
- Experience integrating AI/ML quality checks into CI/CD
- Good understanding of LLM and AI agent architectures
Nice to have
- Hands-on experience with AWS AgentCore Evaluation
- Experience with AWS Bedrock Guardrails, including PII detection
- Knowledge of CloudWatch metrics and production monitoring
- Experience with OpenTelemetry
- Familiarity with LangGraph, Strands Agents, or similar agent frameworks