Project description
Senior data engineer developing and operating Databricks / Spark pipelines and Delta Lake lakehouse layers on Azure for the Client. Accountable for pipeline reliability, data freshness and dataset quality against agreed SLAs. Key tasks • Develop and operate Databricks / Spark pipelines (PySpark, SQL, Delta Live Tables or Workflows). • Design Delta Lake / lakehouse layers (bronze–silver–gold), partitioning and Unity Catalog governance. • Build ETL/ELT jobs and orchestration with Azure Data Factory and/or Airflow; manage dependencies and retries. • Implement data-quality checks and validation (expectations, reconciliation, anomaly alerts). • Tune Spark job performance and cluster cost (autoscaling, Photon, job clusters, spot). • Manage schema evolution and change control; document lineage and transformations.
Responsibilities
SKILLS
Must have
Nice to have
• Databricks Certified Data Engineer Professional; Azure DP-203. • Streaming (Structured Streaming, Kafka / Event Hubs). • dbt, Power BI semantic models, MLflow. • Experience in financial services, sovereign wealth / investment holding or other regulated enterprise environments. • Experience working with distributed teams (onsite UAE with nearshore India / offshore Poland squads).