You'll join Finom’s AI Team as the founding IC dedicated to the quality, evaluation and telemetry of AI agents powering Finom's internal operations and tools (including Ops workflows and internal AI analytics engines across ~20 core processes).
Our belief: an AI agent is only as good as the evaluation loop running on it. Because internal operational agents directly touch financial, compliance and support workflows, evaluation is an exact statistical and engineering discipline here.
Your mission: design the evaluation methodology, build golden benchmarks, and establish quality gates for our internal AI agents from scratch—working directly with process owners and domain experts, with no senior quality owner above you to lean on.
Core Stack: Databricks, DeepEval, Claude Code, Cursor, Python, SQL, dbt.
What You Will Be Doing
Own and extend our offline/online evaluation suites across ~20 internal AI agent processes—datasets (capability + regression), LLM-as-a-judge rubrics, and deterministic checks.
Establish pre-launch quality gates: enforce pass/fail thresholds in CI/CD pipelines before agent prompt, context, or tool changes hit production.
Work directly with domain experts to label cases and resolve annotator disagreement—fixing definition criteria rather than averaging disagreement away.
Build test datasets derived from real user & operational traffic (tickets, internal chats, colleague queries) rather than synthetic edge cases.
Harden statistical methodology: handle judge drift, verbosity bias, non-determinism, and measure true metric shifts vs. noise.
Translate quality numbers into operational decisions: run weekly syncs with process owners to define clear quality vs. cost/latency trade-offs.
Must-Haves
5+ years in Data Science / Product Analytics / Applied AI roles, with sustained product-level metric ownership.
Production LLM Experience: In the last 1–2 years, you have built, shipped, or evaluated LLM-based systems (RAG, multi-step tool use, agents) as a core, primary job responsibility.
Autonomous Quality Ownership: Proven track record of owning evaluation methodology or analytics for an entire product or end-to-end process (what to build vs. what NOT to build).
Fluent Python & SQL: Ability to write clean data pipelines, evaluation harnesses, and dbt transformation models directly.
Statistical Rigor: Applied knowledge of sampling, hypothesis testing, variance analysis, and confidence intervals on noisy metrics.
Daily Setup & Tooling
AI-assisted coding (Claude Code, Cursor, or Codex) is your default daily authoring environment for Python, SQL, and evaluation scripts—not something you occasionally experiment with.
You can walk us through concrete work tasks from the last month where AI coding tools accelerated your engineering and data analysis.
How we work — one thing we mean seriously
AI-assisted coding is our default authoring environment, not a bonus
Claude Code is our main tool — you'll reach for it for SQL, Python, analyses, dashboards, and internal scripts
We're looking for analysts who are already curious and fluent with AI coding — or genuinely excited to become fluent fast
We care about what you ship and how clearly you think
If this idea excites you rather than worries you, you'll feel at home here