Lead Data Software Engineer

EPAM·Удалённо·Удалённо·вчера

We are seeking a hands-on and experienced Lead Data Software Engineer with a deep understanding of Apache Spark and large-scale distributed computing. In this role, you will be responsible for the end-to-end lifecycle of our data infrastructure—from designing and implementing ingestion solutions to ensuring the long-term reliability and scalability of our data pipelines.

The successful candidate will be a collaborative professional, capable of working with cross-functional teams to translate complex business requirements into high-quality technical solutions while maintaining a rigorous focus on engineering standards and performance.

Responsibilities

  • Design, Implementation & Architecture: Lead the design and implementation of robust data ingestion solutions and pipelines using Cloud Native and Big Data technologies. You will be responsible for architecting key system components and selecting the appropriate tech stack to ensure long-term scalability, performance, and alignment with the broader technical roadmap
  • Team Leadership & Customer Communication: Act as the team lead by providing daily guidance and mentorship to ensure the team delivers high-quality, scalable solutions on schedule. Serve as the primary technical point of contact for customers, translating business visions into actionable technical roadmaps while managing expectations regarding delivery, feasibility, and progress
  • System Management: Manage the full technical stack, including configuration management, monitoring, debugging, and performance tuning of data solutions
  • Pipeline Maintenance: Develop and maintain scalable data pipelines for efficient data processing and analysis
  • Collaboration: Partner with cross-functional teams and data architects to identify business requirements and translate them into technical specifications
  • Quality Assurance: Lead and participate in code reviews and development testing to ensure adherence to standards, quality gates, and SDLC best practices
  • Documentation: Create and maintain detailed technical documentation for all data engineering projects to ensure transparency and knowledge sharing
  • Operational Excellence: Troubleshoot and resolve data-related issues in a timely and efficient manner, ensuring all solutions are of high quality, reliable, and scalable

Requirements

  • Professional Experience: 5+ years of experience in Data Software Engineering focused on large-scale distributed systems
  • Architectural Design & Scalability: Demonstrated track record of designing modular system components and selecting technology stacks (e.g., storage formats, processing engines) that ensure long-term performance, scalability, and alignment with the broader technical roadmap
  • Technical Leadership & Customer Communication: Experience as a Team Lead providing technical guidance to the team to ensure high-quality delivery, while serving as a technical voice for customers to help bridge the gap between business needs and engineering solutions
  • Programming Languages: Proficient in Python and SQL
  • Big Data Stack: Extensive hands-on experience with Apache Spark, Kafka, and Airflow
  • Cloud Infrastructure: Proficient with major cloud providers (AWS, Azure, or GCP)
  • Platform Knowledge: Databricks or Snowflake (nice to have)
  • Technical Proficiency: Practical experience in Apache Spark job performance tuning
  • Engineering Standards: Strong understanding of the SDLC, quality gates, and Agile methodologies
  • Soft Skills: Excellent communication skills with English proficiency at a B2+ level
  • Attributes: Motivated, independent, and capable of handling several projects simultaneously with a focus on efficiency

Похожие вакансии

Другие вакансии EPAM