We are seeking a Lead AI Engineer to build and evolve the software behind a line's AI systems and lead their development toward agentic systems — building the services and workflows that LLM-based agents use. You'll work day-to-day with agentic coding harnesses (e.g., Claude Code): break work into well-scoped tasks, direct AI agents to build and modify code, and review the output with calibrated scrutiny — spec-first, tested, and accountable for what your agents produce.
Responsibilities
- Maintain, refactor, and extend the line's AI applications and services in production
- Build agentic-friendly software: the tools, services, and workflows that LLM-based agents call
- Design AI workflows that combine models, prompts, tools, enterprise data, and business logic
- Validate model behavior, outputs, and assumptions from an engineering and production perspective
- Design and maintain evaluation frameworks for AI/LLM systems, including test suites and human evaluation workflows
- Write evaluations and quality checks for agent behavior; build in grounding and assurance
- Ensure reliability, scalability, cost control, and latency optimization
- Help bring new agentic systems into production
Requirements
- 5+ years of experience in AI Solution Engineering
- At least 1 year of relevant leadership experience
- Strong software engineering proficiency in Python as the primary language, with other modern languages such as TypeScript, Java/Kotlin acceptable where production-quality code delivery is demonstrated
- Forward Deployed Engineer mindset: desire to understand the business problem, the impact the solution makes, and drive pragmatic decisions toward realising it — including API, service, and integration engineering
- Hands-on experience with LLM application patterns (prompting, retrieval-augmented generation, tool/function calling) and proven agentic development practice
- Expert daily practitioner competency in agentic coding harnesses (e.g., Claude Code): spec-first decomposition, directing agents to build and modify production code, critically reviewing and owning generated output
- Capability to pick up and improve an existing codebase
- Skills in validating AI model behavior and outputs from an engineering and production standpoint
- Background in hands-on agentic/LLM application development: agent loops, evaluations, and tool design
- English proficiency at an Upper-Intermediate level (B2) or higher
Nice to have
- Familiarity with Palantir Foundry / AIP
- Insurance or reinsurance domain exposure
- Knowledge of cloud AI services (e.g., Azure AI)