The Role
We are building the AI layer of our sales organisation: an agentic system used daily by ~190 relationship managers across 14 regions, working against our CRM, Jira and internal knowledge base. It answers questions on clients and portfolios, delivers briefs, manages tasks and suggests the next best action.
Next up: real coaching against a manager's targets and sales methodology, AI role play, proactive task suggestion driven by data signals, deeper integrations, and the evaluation infrastructure that tells us any of it works.
Small team, high autonomy, no groomed board. You take a scope, build it end to end and own how it behaves in production.
Stack: TypeScript / Node 22, Express 5, Drizzle ORM + PostgreSQL, Redis, Slack Bolt, Vertex AI (Gemini), Langfuse, Vitest, GitLab CI, Docker. React 19 + Vite + Tailwind/shadcn on the frontend. Python for evals, data work and glue. Roughly 80/20 backend to frontend — and you are not expected to know all of it, you are expected to get up to speed on your own.
What you'll do
• Build and own the agent layer: harness design, prompt architecture, the tool layer, orchestration and context strategy.
• Build and own evaluations — the mechanism that tells us a change made the agent better rather than merely different.
• Build the backend behind it: APIs, services, data models, integrations with company systems — plus the frontend that exposes them.
• Design with an attacker in mind: prompt injection, cross-tenant data leakage, access control. We handle client financial data.
• Extend the scope where it needs it — edge cases, failure modes, UX, and the metric that proves it worked.
• Dig into unfamiliar infrastructure, find the right people, get to a decision.
What you'll bring
• Agent engineering — you have built LLM agents in production, not demos: harness structure, prompt architecture, tool calling, RAG, MCP, context management, and debugging pipelines that fail differently every time. This is the core of the role.
• Evaluation discipline — you treat "did this change help?" as an engineering question with an answer, not a vibe check.
• Strong backend engineering in TypeScript/Node, plus working Python. You own your features all the way to the screen, so you need to build a competent frontend without waiting for someone else.
• Fluency building with agents — you use Claude Code or an equivalent harness as a real force multiplier: skills, subagents, hooks, custom tooling. And what comes out the other end is a deliberate, secure, maintainable feature, not the volume of half-understood code these tools make it so easy to generate. You own every line you ship, whether or not you typed it.
• Autonomy — you take a scope rather than instructions and carry it to done, including the parts nobody wrote down.
• Security instinct — the habit of asking "how would someone break this" before shipping.
• Taste in interfaces, and we mean something specific: sufficient minimalism. In the age of AI it is trivially easy to build a Christmas tree that looks impressive and solves nothing.
• Working proficiency in Russian and English. Our engineering teams work in Russian; documentation and company-wide collaboration are in English.
Nice to have: a background in financial services, brokerage or fintech; Slack apps, CRM integrations, OAuth / on-behalf-of flows; agent observability and tracing at scale.
How you work matters as much as what you know. You propose solutions instead of raising problems. You start simple and earn complexity — a thin working version in front of real users beats an elegant one that ships next quarter. And you think about what happens after the merge: who uses this, does it make their job faster, and how will we know.
What we offer