About the Position
We are looking for a Senior ML Evaluation Engineer to help design, implement, and operationalize evaluation frameworks for enterprise AI systems. In this role, you will define quality standards for AI agents and machine learning solutions, build scalable evaluation pipelines, and integrate automated quality gates into CI/CD processes. You will work closely with AI engineers, platform teams, and governance stakeholders to ensure reliable, measurable, and production ready AI outcomes.
About the Project
The project focuses on establishing enterprise grade evaluation standards for AI agents and machine learning systems. The platform enables automated quality assessment, production monitoring, and governance through evaluation frameworks, observability data, and continuous validation processes.
About the Team
You will work within a collaborative team of ML engineers, AI platform engineers, software developers, architects, and product stakeholders. The team follows a data driven approach to AI quality, emphasizing automation, observability, reliability, and continuous improvement.
Responsibilities
- Develop and maintain evaluation frameworks for AI agents and machine learning solutions
- Design LLM as judge evaluation methodologies using built in helpfulness and correctness evaluators
- Build custom Python based evaluators to perform deterministic quality and compliance checks
- Define enterprise evaluation standards, including mandatory assessment dimensions and pass or fail criteria
- Implement evaluation workflows across response, tool invocation, and end to end session levels
- Integrate observability telemetry and OpenTelemetry spans into evaluation pipelines
- Design and maintain CI/CD quality gates for machine learning models and AI agents
- Collaborate with AI platform teams to improve evaluation coverage, automation, and reporting
- Analyze evaluation results and provide recommendations to improve agent reliability and performance
- Support production monitoring strategies and continuous quality verification processes
- Contribute to AI governance initiatives and best practices for model and agent evaluation
Requirements
- 5+ years of experience in machine learning engineering or AI platform engineering
- Hands on experience designing and implementing LLM evaluation frameworks
- Experience creating custom evaluators for deterministic quality validation and policy enforcement
- Experience building CI/CD deployment gates for machine learning models, AI applications, or agent based systems
- Strong Python development skills
- Experience working with AI quality metrics, automated testing methodologies, and evaluation pipelines
- Understanding of agent based architectures and modern AI application development practices
- Experience collaborating with engineering teams on quality assurance and governance initiatives
- Strong analytical and problem solving skills
- Effective written and verbal communication skills
Nice to Have
- Hands on experience with AWS Agent Evaluation APIs, including evaluation execution and results analysis
- Experience integrating AWS Bedrock Guardrails for PII detection and evaluation workflows
- Experience using CloudWatch metrics for online evaluation monitoring and reporting
- Knowledge of observability frameworks and OpenTelemetry based monitoring
- Experience with enterprise AI governance and compliance programs
- Experience evaluating production AI agents and large scale machine learning systems
Technologies
AWS Agent Evaluation Services, Python, AWS Lambda, OpenTelemetry, CI/CD Pipelines, Machine Learning Evaluation Frameworks, Large Language Models, Evaluation APIs, CloudWatch, AWS Bedrock Guardrails, Enterprise AI Governance Tools