Agent Evaluation Engineer — Build-Time Framework & Deployment Gates

EPAM·Portugal·Удалённо·сегодня

We're looking for an Agent Evaluation Engineer — Build-Time Framework & Deployment Gates to join our team in Portugal in a fully remote working mode. In this role, you will design and maintain an evaluation framework for AI agents, ensuring quality and compliance through automated tests and CI/CD deployment gates. You will develop multi-layer evaluation suites that blend deterministic checks with LLM-powered graders, simulate multi-turn conversations, and define reliability metrics. The position also involves implementing staging validations, shadow-mode traffic analysis, and A/B rollout strategies, with feedback loops from production environments to enhance overall system robustness.

Responsibilities

  • Design and implement build-time evaluation frameworks for agentic workflows using LangGraph or comparable orchestration frameworks
  • Create deterministic and LLM-as-judge grading pipelines covering reasoning, trajectory accuracy, and output quality
  • Develop test harnesses for multi-turn conversational simulations and context-retention scoring
  • Define reliability assessment methods including multi-trial metrics (pass@k, pass^k)
  • Implement CI/CD deployment gates that enforce quality thresholds and block releases not meeting standards
  • Integrate staging validation, shadow-mode traffic comparison, and A/B rollout control in deployment pipelines
  • Leverage AWS AgentCore Evaluations for on-demand and online scoring components connected to production feedback
  • Convert production incidents into reusable regression cases for continuous quality improvement
  • Collaborate with engineering and DevOps teams to embed evaluation gates into automated workflows

Requirements

  • 4+ years of experience in building automated testing or evaluation frameworks for ML, LLM, or agentic systems
  • Proven hands-on experience designing multi-layer evaluation suites with deterministic and LLM-based graders
  • Expertise with CI/CD pipelines and implementing metric-based quality gates for automated deployments
  • Practical knowledge of LangGraph or similar agent orchestration frameworks
  • Strong background in designing simulation-based evaluation strategies and conversation-level tests

Nice to have

  • Experience with AWS AgentCore Evaluations API (CreateEvaluation, custom evaluators)
  • Familiarity with shadow-mode, canary, or A/B deployment practices for ML-based platforms
  • Background in transforming production failures into build-time regression tests for agent workflows

Похожие вакансии

Другие вакансии EPAM