As the Lead AI Engineer for our next-generation TCO Agent Platform, you will serve as the technical lead and multi-agent system architect for a greenfield, proactive FinOps AI platform. You will oversee the development, orchestration, and governance of an ecosystem comprising 8 specialized AI agents (e.g., Commitment Optimization, Financial Operations, Resource Optimization, Anomaly Detection) running within a multi-cloud GCP/AWS environment.
You will drive core architectural execution, ensure strict adherence to SOX-adjacent financial controls, enforce Zero Trust security models via Model Armor and LiteLLM, and guide the engineering team through building high-scale, autonomous enterprise AI services.
Reporting directly to the Technical Product Manager, you will collaborate closely with Solution, Platform, and Enterprise Architects. You will also have a Senior Software Engineer under your direct subordination to collaborate with on solution implementation.
Req.#1087619193
Responsibilities
- Multi-Agent Architecture & Orchestration: Lead the design and implementation of 8 specialized agents using Python, FastMCP, and GCP Workload Identity. Oversee inter-agent dependencies, prompt engineering lifecycle, tool definitions, and agent-to-service communication
- LLM Governance & Tokenomics: Enforce centralized LLM routing via LiteLLM Gateway and Vertex AI Model Garden (Claude, Gemini). Implement agent self-governance tracking systems (Tokenomics) to monitor and cap LLM operating costs within strict platform operational limits
- Financial & Compliance Guardrails: Architect execution boundaries and strict segregation-of-duties workflows for SOX-adjacent processes (e.g., Journal Entry generation vs. human approval)
- System Integration & Action Routing: Oversee the architecture of the platform’s Action Gateway—handling direct API invocations, event-driven workflows, and fallback ticketing (Jira, Slack, Teams)
- Technical Leadership & Standards: Set coding, testing, and formatting standards across application repositories. Mentor Senior and Mid-level AI engineers and drive code reviews enforcing RFC standard error formats, API contracts, and schema compliance
Requirements
- Experience: 8+ years of software engineering experience with 3+ years in a technical leadership capacity building multi-agent AI systems, FinOps tools, or LLM-powered platforms
- Frameworks & Languages: Advanced proficiency in Python 3.11+, FastMCP, FastAPI, and Pydantic. Prior experience with object-oriented enterprise languages (e.g., Java) for seamless integration with core platform services and backend APIs
- AI/LLM Architecture: Hands-on experience with Google ADK, Vertex AI, LiteLLM Gateway, Model Armor guardrails, prompt engineering, structured tool output parsing, and agent execution boundaries
- Data & Cloud Platforms: Deep familiarity with the GCP ecosystem (BigQuery, GKE, Workload Identity), SQL schema design (FOCUS standard preferred), and partitioned/clustered OLAP architectures
- Security & Governance: Experience implementing Zero Trust authentication (OAuth/KSA-to-GSA), RBAC, immutable audit logging, and API/MCP error specifications (RFC 7807/9457)
- DevOps & Infrastructure: Proficiency with Docker builds, GKE deployment patterns, OpenTofu/Terraform, and CI/CD pipelines
- Leadership Skills: Proven ability to coach, mentor, and influence teams beyond just writing and implementing solutions
Nice to have
- Experience with LangGraph / LangChain
- Hands-on experience with Vertex AI
- Working knowledge of modern DevOps and CI/CD practices
- Familiarity with GCP infrastructure resources and their constraints, with the ability to identify optimal resources for designed AI agentic solutions