We are looking for a Senior Data Software Engineer to join the SAP Data Engineering CoE, building and scaling data pipelines in Databricks with PySpark and Python to deliver data products for ML modeling teams.
Responsibilities
Support a team of data engineers to build pipelines used by MLOps and ML Engineers on ML modeling teams
Develop, optimize, and maintain data transformation pipelines in Databricks using PySpark
Work with data stored in ADLS Gen2 and SAP HANA Data Lake, primarily in Delta/Parquet format
Implement and maintain data quality checks, including schema validation, deduplication, enrichment, and tagging
Communicate with stakeholders to understand business processes and model input data
Tune performance for large-scale datasets to ensure efficient processing
Requirements
3+ years of experience in data engineering with proficiency in Python, PySpark, and Databricks
Hands-on experience with Delta Lake and Azure data lake technologies
Familiarity with software version control tools such as GitHub and Git
Experience with CI/CD frameworks such as GitHub Actions
Knowledge of data lake technologies and large-scale dataset performance tuning
Proficiency in English (B2+) for effective communication with the customer's team
Nice to have
Familiarity with at least one other programming/scripting language, such as Java, SQL, or Scala