Our team is seeking a dynamic, highly experienced professional to fill the role of Lead Operational Intelligence Engineer.
This role involves taking charge of developing, maintaining, and enhancing our cloud-based Elastic & Observability Platform. The successful candidate will spearhead strategic initiatives, mentor a top-performing technical team, and maintain platform reliability while promoting innovation and self-service capabilities for platform users. On-call rotation duties for monitoring platform health and functionality are also part of this position.
Responsibilities
- Ensure observability and search platforms exceed business SLAs in terms of availability, functionality, performance, and security
- Deliver technical leadership when complex incidents arise and ensure prompt escalation of resolutions during on-call shifts
- Create and maintain thorough platform documentation, standard operating procedures, and knowledge-sharing materials
- Work with cross-functional teams, stakeholders, and vendors to manage operational needs, advance strategic initiatives, and handle installations, troubleshooting, and upgrades
- Drive improvements to platform features and self-service tools, including advanced Elastic Synthetics and automated chargeback processes
- Design and build proofs-of-concept to advance platform innovation, such as AI-driven observability, sophisticated data processing models, or migration to Kubernetes-based platforms
- Guide the construction, deployment, and upkeep of Elastic clusters using Infrastructure-as-Code tools such as Terraform and Ansible, and coach team members on best practices
- Manage platform lifecycle tasks, such as component upgrades, capacity planning, cost optimization, and adapting to new compliance requirements
- Regularly evaluate and optimize ELK stack performance, covering ingestion, indexing, and query tuning for large-scale environments
- Build and improve alerting and incident management processes by integrating advanced monitoring tools like Kibana Rules, Watchers, and PagerDuty
- Manage the ingestion, enrichment, backup, and restoration of large-scale platform data, optimizing data workflows along the way
- Direct and plan major operational events, including SSL certificate rotations, cluster migrations, and scalability optimization efforts
Requirements
- At least 5 years of experience in Operational Intelligence, demonstrating leadership and technical skill in managing large-scale observability platforms
- Proven ability to design and oversee Elastic clusters within complex, multi-cloud environments
- Comprehensive knowledge of Elastic Stack components, including advanced setups of Elasticsearch, Kibana, and Logstash
- High-level skills in Infrastructure-as-Code tools such as Terraform and Ansible, with flexibility to work with tools like Jenkins CI or GitOps frameworks
- Strong Python scripting abilities for automating processes, handling data, and expanding platform interoperability
- Solid grasp of incident management frameworks and workflows using tools such as PagerDuty, Uptrends, and other enterprise monitoring platforms
- Demonstrated success in diagnosing and resolving intricate platform issues within strict SLA timeframes
- Strong skills in managing and scaling fault-tolerant platforms, ensuring performance, security, and compliance across large distributed systems
- Proven track record of mentoring team members, managing priorities, and serving as a liaison between technical and non-technical groups
- Strong English communication skills (B2+ level), both written and verbal, with an emphasis on technical communication
Nice to have
- Skills in Groovy scripting or advanced Linux administration experience to streamline platform operations
- History of enhancing observability workflows through custom integrations in tools such as Uptrends, PagerDuty, or Elastic
- Practical experience configuring advanced Elastic Synthetics for reliable monitoring and custom synthetic testing
- Background in leading strategic initiatives like AI-driven modernization, cloud-native migrations, or cost-saving observability improvements