Python Developer with AI Evaluation Frameworks

DataArt·Armenia, Bulgaria, Cyprus, Georgia, Kazakhstan, Latvia, Poland, Romania, Serbia, Ukraine·Удалённо, Офис·сегодня

About the Position

We are looking for a Python Developer with experience in AI evaluation and quality assurance to support the development of an enterprise AI quality platform. You will design and implement automated evaluation frameworks that measure the quality, safety, accuracy, and compliance of AI generated outputs. The role involves building custom evaluators, integrating cloud based AI services, and supporting governance and monitoring initiatives across AI solutions.

About the Project

The project focuses on building a centralized AI evaluation platform that enables automated validation of AI applications, workflows, and agent interactions. The platform provides quality controls for model responses, workflow compliance, PII detection, tool validation, and operational monitoring to support enterprise AI adoption.

About the Team

You will work within a cross functional team of AI engineers, backend developers, platform specialists, architects, and governance stakeholders. The team follows an iterative delivery model with a strong emphasis on quality, security, observability, and continuous improvement.

Responsibilities

  • Develop custom AI evaluation services using Python and AWS Lambda
  • Build LLM as judge evaluators to assess subjective quality dimensions of AI generated content
  • Implement tool response validation using JSON Schema based evaluation mechanisms
  • Develop workflow contract compliance checks at session level
  • Create numerical accuracy validation logic and trace level evaluation controls
  • Implement PII detection mechanisms using pattern matching techniques and AWS Bedrock Guardrails
  • Integrate evaluation services with enterprise AI platforms and workflows
  • Design and maintain reusable evaluation frameworks and quality standards
  • Collaborate with AI engineers and architects to define evaluation methodologies
  • Document evaluation rules, quality metrics, and implementation standards
  • Support monitoring, troubleshooting, and continuous improvement of evaluation services

Requirements

  • 3+ years of experience in Python software engineering
  • Experience implementing quality assurance, validation, or evaluation frameworks for AI or machine learning systems
  • Experience developing and deploying AWS Lambda functions
  • Experience designing backend services and data validation processes
  • Knowledge of prompt engineering techniques for AI assessment and evaluation
  • Experience working with JSON based APIs and schema validation approaches
  • Understanding of AI quality metrics, testing methodologies, and governance concepts
  • Experience working in cloud environments and modern software development practices
  • Strong analytical and problem solving skills

Nice to Have

  • Experience registering custom evaluators within AWS AgentCore Evaluation
  • Experience integrating AWS Bedrock Guardrails for PII detection and content validation
  • Experience using CloudWatch Logs for monitoring, reporting, and evaluation result tracking
  • Experience with AI observability and monitoring frameworks
  • Knowledge of enterprise AI governance and compliance requirements
  • Experience building automated quality control solutions for AI platforms
  • Familiarity with large language model evaluation methodologies and benchmarking techniques
  • Experience working with enterprise scale AI workloads and cloud native architectures

Technologies

Python, AWS Lambda, AWS AgentCore Evaluation, AWS Bedrock, AWS Bedrock Guardrails, CloudWatch Logs, JSON Schema, pytest, AWS

Похожие вакансии

Другие вакансии DataArt