Senior AI Quality Engineer (Hyderabad-based)

DataArt·India·Офис·3д. назад

About the Position

We are looking for a Senior AI Quality Engineer who can design the quality strategy and still write the tests. You will build deterministic checks and probabilistic evaluations, automate UI and API workflows, test data pipelines, and turn model, data, security, and operational risks into release gates.  This position is intended for Hyderabad-based candidates who are available to work from the office.

About the Project

Our client, a globally recognized leader in investment management education and certification, ships AI agents, retrieval systems, and data pipelines. Quality for that work is part of delivery: evaluation gates, automated regression, security and privacy checks, and evidence that a release is ready.

Responsibilities

  • Define AI evaluation strategies, golden datasets, thresholds, and release gates. Automate response, retrieval, trajectory, and contract checks, calibrate LLM-as-judge, account for non-determinism and drift, and classify failures by consequence.
  • Write and debug Python automation for AI, API, data, and integration workflows, and TypeScript and Playwright automation for UI and API workflows. Use AI-assisted tools, including agents, and review generated tests for missing coverage.
  • Integrate AI evaluations and UI/API regression suites into CI/CD, interpret Allure or equivalent reports, and keep TestRail or equivalent aligned with automated coverage and CI results.
  • Test REST and streaming APIs, agent tool contracts, and data and retrieval pipelines, including authentication, sessions, errors, compatibility, call budgets, idempotency, lineage, and rollback.
  • Test prompt injection, jailbreaks, unsafe tool use, data exfiltration, access control, tenant and role isolation, PII controls, audit events, guardrails, and human approval for consequential actions.
  • Measure latency, throughput, token use, and cost. Verify timeout, retry, fallback, and loop-control behavior, diagnose failures across services, and produce release and post-deployment evidence.
  • Work with domain experts, product, security, engineering, and operations on ground truth, traceable coverage, and untested scope.

Requirements

  • Hands-on AI evaluation experience: strategies, golden datasets, LLM-as-judge calibration, non-determinism, drift, failure severity, and release gates.
  • Proficiency in Python and pytest, or an equivalent framework, and strong hands-on TypeScript and Playwright experience for UI and API automation.
  • Experience using AI-assisted development tools, including agents, and reviewing generated tests for missing coverage.
  • Experience integrating evaluation and UI/API regression suites into CI/CD, using test reporting, and keeping test-case management aligned with automation and CI results.
  • Experience testing REST and streaming APIs, agent tool contracts, and data or retrieval pipelines.
  • Experience with adversarial, access-control, privacy, and guardrail testing, plus performance, cost, and distributed failure diagnosis.
  • Experience producing release evidence and working with domain experts and delivery partners on ground truth and coverage gaps.

Nice to Have

  • Experience selecting quality metrics by system type and producing audit-ready evaluation reports.
  • Shared automation framework structures, reusable components, code reviews, Git and GitHub workflows, test data management, and test-environment troubleshooting.
  • Cloud cross-browser grids such as LambdaTest or equivalent.
  • Security finding triage, operational readiness checks, and improvement of shared evaluation patterns, automation libraries, and datasets.

Technologies

Python, pytest or an equivalent test framework, promptfoo or an equivalent AI evaluation harness, TypeScript, Playwright, REST and SSE API automation, Git and GitHub, GitHub Actions or an equivalent CI/CD platform, TestRail or an equivalent test-management tool, Allure Report or an equivalent reporting tool, AI-assisted development tools (Cursor, GitHub Copilot, or equivalent), including agents

Похожие вакансии

Другие вакансии DataArt