We are looking for a Lead Data DevOps Engineer to join our team. We are building an Enterprise AI Gateway from the ground up. It serves as the single entry point every team in the company uses to access large language models. It is also how we maintain control over AI at scale: who can use which models, what it costs, and what gets logged. The goal is for it to be self-sustaining. Onboarding teams, agents, and MCP servers, along with managing keys, permissions, limits, and guardrails, should happen through automation rather than manual tickets. Very little of this exists yet, so you will influence it starting from the earliest design choices. The team is small and focused, meaning decisions move quickly and your contributions will be clearly visible. We are looking for someone who defaults to automation and who values the governance side of AI just as much as the models themselves.
Responsibilities
- Architect and construct the foundational design of the Enterprise AI Gateway starting from scratch
- Automate onboarding processes for teams, agents, and MCP servers, including the provisioning of keys, permissions, limits, and guardrails
- Build and maintain infrastructure as code to enable scalable, self-service platform capabilities
- Set up monitoring, logging, and cost-tracking systems to preserve visibility and oversight of AI usage
- Establish guardrails and governance mechanisms that control which teams and models can access particular resources
- Guarantee the reliability, uptime, and performance of production systems that power the gateway
- Work directly with platform users, resolving questions and troubleshooting errors as they arise
- Refine automation on an ongoing basis to minimize manual work and reliance on ticket-based workflows
- Assess and incorporate emerging GenAI and agentic AI patterns, frameworks, and protocols into the platform
- Play a role in shaping major architectural and design decisions as the platform matures from its early stages
Requirements
- At least 5 years of relevant experience
- A minimum of one year of experience leading and managing teams
- Solid Python skills for automation, extensions, and integrations
- Strong SRE capabilities, backed by genuine experience maintaining healthy production systems
- Proficient with Git for version control
- Experience working with Google Cloud Platform
- Familiarity with LLMOps practices
- Experience using Terraform and Helm for infrastructure automation
- Working knowledge of GenAI/Agentic AI concepts, including associated patterns, frameworks, and protocols
- Strong communication skills, with the ability to clearly convey technical concepts, as you will engage daily with platform users, answering questions and resolving issues
- Excellent English proficiency (B2 level or higher)
Nice to have
- Hands-on experience with Google Vertex AI, especially endpoints for model serving and Model Armor
- GCP experience involving services such as BigQuery, Cloud Run, and IAM
- Experience building AI agents, for instance using Google's Agent Development Kit (ADK)
- Experience with AWS Bedrock
- Extensive Kubernetes experience, ideally with GKE
- Familiarity with Groovy
- Experience with CI/CD pipelines using Jenkins
- Production experience with an AI gateway, such as LiteLLM or EPAM DIAL, considered highly valuable