We are building reusable, governed data-sharing adapters that connect a cloud lakehouse and analytics warehouse to external platforms with strict access controls and SLAs. As a Lead Data Software Engineer, you will define lakehouse patterns, deliver integrations, and reinforce lineage and governance across data-out flows.
Responsibilities
- Design a UniForm write layer using one physical dataset with dual-format metadata for Delta and Iceberg consumers
- Build and validate GCS-to-BigQuery ingestion pipeline patterns for structured operational data
- Implement Kafka-based CDC patterns for real-time and near-real-time movement into the lakehouse
- Develop dependency-aware bookkeeping plus data lineage tracking patterns across pipelines
- Engineer modular, version-controlled adapter code intended for reuse across new integrations
- Configure Iceberg external table definitions in Snowflake using Horizon Catalog governance features
- Validate zero-copy read access from Snowflake to Iceberg and Delta tables without data movement
- Implement tenant-scoped access controls aligned with Snowflake metadata governance requirements
- Deliver and certify a Delta Sharing adapter for live, zero-copy sharing from Delta Lake to Databricks consumers
- Configure Delta Sharing endpoints and manage sharing agreements for Databricks access
- Validate Databricks read access via Delta Sharing for Spark, Pandas, and compatible consumers
- Test end-to-end freshness and sharing latency to meet agreed SLA targets
- Register connector types and implement RBAC plus tenant-scoped authorization for all data-out paths
- Implement metering hooks compatible with billing requirements for governed data-out flows
Requirements
- Proven experience with 5+ years in data engineering using Python and cloud data platforms
- Hands-on expertise with Google Cloud BigQuery for advanced analytics and data modeling
- Solid background working with lakehouse table formats, including Apache Iceberg and Delta Lake
- Demonstrated leadership skills to steer integration design, technical decisions, and delivery ownership
- Strong track record executing projects that deliver data pipelines and adapters meeting SLAs
- Advanced skills in Kafka/CDC patterns for real-time and near-real-time data movement
- Deep understanding of data lake and lakehouse ingestion architecture on object storage
- Excellent collaboration skills to align with platform, governance, and downstream consumer needs
- Upper-Intermediate English proficiency (B2) for technical discussions and documentation
Nice to have
- Apache Spark experience for validation and consumption testing
- Knowledge of Databricks Unity Catalog for governed access patterns
- Delta Lake expertise, including Delta Sharing setup and troubleshooting
- Gen AI Assisted Development experience with tools such as Claude Code, GitHub Copilot, or Cursor
- Snowflake Horizon Catalog experience for metadata governance and external table management