We're looking for a hands-on, experienced Lead Data Software Engineer who has a deep understanding of Apache Spark and large-scale distributed computing. This role covers the end-to-end lifecycle of our data infrastructure, from designing and implementing ingestion solutions to ensuring the long-term reliability and scalability of our data pipelines.
The ideal candidate is a collaborative professional who can work with cross-functional teams, translate complex business requirements into high-quality technical solutions, and maintain a rigorous focus on engineering standards and performance.
Responsibilities
- Design and implementation of robust data ingestion solutions and pipelines using Cloud Native and Big Data technologies, including architecting key system components and selecting the appropriate tech stack to ensure long-term scalability, performance, and alignment with the broader technical roadmap
- Team leadership through daily guidance and mentorship so the team delivers high-quality, scalable solutions on schedule, along with serving as the primary technical point of contact for customers, translating business visions into actionable technical roadmaps and managing expectations around delivery, feasibility, and progress
- Full technical stack management, including configuration management, monitoring, debugging, and performance tuning of data solutions
- Development and maintenance of scalable data pipelines for efficient data processing and analysis
- Partnership with cross-functional teams and data architects to identify business requirements and translate them into technical specifications
- Code reviews and development testing leadership and participation to ensure adherence to standards, quality gates, and SDLC best practices
- Creation and maintenance of detailed technical documentation for all data engineering projects to support transparency and knowledge sharing
- Timely and efficient troubleshooting and resolution of data-related issues, ensuring all solutions remain high-quality, reliable, and scalable
Requirements
- 5+ years of experience in Data Software Engineering with a focus on large-scale distributed systems
- A demonstrated track record of designing modular system components and selecting technology stacks (such as storage formats and processing engines) that ensure long-term performance, scalability, and alignment with the broader technical roadmap
- Experience as a Team Lead providing technical guidance to ensure high-quality delivery, while acting as a technical voice for customers to bridge the gap between business needs and engineering solutions
- Proficiency in Python and SQL
- Extensive hands-on experience with Apache Spark, Kafka, and Airflow
- Proficiency with major cloud providers (AWS, Azure, or GCP)
- Knowledge of Databricks or Snowflake (nice to have)
- Practical experience in Apache Spark job performance tuning
- Strong understanding of the SDLC, quality gates, and Agile methodologies
- Excellent communication skills with English proficiency at a B2+ level
- Motivation, independence, and the ability to handle several projects simultaneously with a focus on efficiency