Senior AI Evaluation Engineer with LangGraph

DataArt·Armenia, Bulgaria, Cyprus, Georgia, Kazakhstan, Latvia, Poland, Romania, Serbia, Ukraine·Удалённо, Офис·вчера

About the Position

We are looking for a Senior AI Evaluation Engineer to design and implement enterprise grade evaluation frameworks for agent based AI systems. In this role, you will build automated quality assessment capabilities, define deployment gate strategies, and create scalable evaluation methodologies that improve the reliability, accuracy, and safety of AI driven solutions throughout the development lifecycle.

About the Project

The project focuses on establishing a comprehensive evaluation and quality assurance platform for agent based AI systems. The solution provides automated testing, deployment validation, production feedback integration, and continuous quality monitoring to ensure high standards for agent performance and user experience.

About the Team

You will work with AI engineers, ML engineers, platform engineers, software developers, and product stakeholders in a collaborative environment focused on quality, reliability, observability, and continuous improvement. The team is responsible for defining evaluation standards and operationalizing AI quality across enterprise platforms.

Responsibilities

  • Design and develop build time evaluation frameworks for LangGraph based agent systems
  • Create automated test harnesses for graph level and node level validation
  • Design evaluation strategies that combine deterministic grading and LLM as judge methodologies
  • Implement multi layer evaluation frameworks covering tool selection accuracy, execution trajectory quality, reasoning effectiveness, and output quality
  • Develop multi turn conversation simulations and context retention scoring mechanisms
  • Implement multi trial reliability testing methodologies, including pass at k and pass power k approaches
  • Design and maintain CI/CD deployment gates that validate quality metrics before production releases
  • Build staging validation, shadow mode comparison, and controlled rollout evaluation workflows
  • Integrate production evaluation feedback into build time testing frameworks to improve quality coverage
  • Collaborate with platform and engineering teams to establish evaluation standards, thresholds, and governance practices
  • Analyze evaluation results and provide recommendations for improving agent reliability and performance
  • Contribute to technical documentation, testing standards, and quality engineering best practices

Requirements

  • 4+ years of experience building automated testing frameworks, evaluation platforms, or quality assurance solutions for machine learning, large language model, or agent based systems
  • Hands on experience designing multi layer evaluation frameworks that combine deterministic and LLM as judge grading approaches
  • Experience implementing automated quality gates that can block deployments based on predefined metric thresholds
  • Experience working with LangGraph or a comparable agent orchestration framework
  • Strong understanding of agent behavior evaluation, workflow validation, and AI quality measurement techniques
  • Experience designing scalable testing and validation processes for production AI systems
  • Strong Python development experience
  • Knowledge of CI/CD practices, deployment automation, and release governance
  • Experience analyzing evaluation data and translating findings into platform improvements
  • Strong communication and collaboration skills

Nice to Have

  • Hands on experience with AWS AgentCore Evaluations, including evaluation execution and custom evaluator development
  • Experience designing and operating shadow mode or canary deployment strategies for machine learning or AI systems
  • Experience creating automated feedback loops that convert production incidents into regression test scenarios
  • Knowledge of production observability, monitoring, and evaluation pipelines
  • Experience with enterprise AI governance and quality assurance programs
  • Understanding of agent observability and telemetry driven quality improvement processes

Technologies

LangGraph, AWS AgentCore Evaluations, Python, LLM as Judge Frameworks, CI/CD Pipelines, OpenTelemetry, Agent Based Systems, Automated Evaluation Frameworks, Large Language Models, A/B Testing, Shadow Mode Validation, Deployment Automation, Quality Monitoring Tools

Похожие вакансии

Другие вакансии DataArt