Lead Data Software Engineer

EPAM·Argentina, Colombia, Mexico, Brazil, Chile·Удалённо·вчера

We are building reusable, governed data-sharing adapters that connect a cloud lakehouse and analytics warehouse to external platforms with strict access controls and SLAs. As a Lead Data Software Engineer, you will define lakehouse patterns, deliver integrations, and reinforce lineage and governance across data-out flows.

Responsibilities

  • Design a UniForm write layer using one physical dataset with dual-format metadata for Delta and Iceberg consumers
  • Build and validate GCS-to-BigQuery ingestion pipeline patterns for structured operational data
  • Implement Kafka-based CDC patterns for real-time and near-real-time movement into the lakehouse
  • Develop dependency-aware bookkeeping plus data lineage tracking patterns across pipelines
  • Engineer modular, version-controlled adapter code intended for reuse across new integrations
  • Configure Iceberg external table definitions in Snowflake using Horizon Catalog governance features
  • Validate zero-copy read access from Snowflake to Iceberg and Delta tables without data movement
  • Implement tenant-scoped access controls aligned with Snowflake metadata governance requirements
  • Deliver and certify a Delta Sharing adapter for live, zero-copy sharing from Delta Lake to Databricks consumers
  • Configure Delta Sharing endpoints and manage sharing agreements for Databricks access
  • Validate Databricks read access via Delta Sharing for Spark, Pandas, and compatible consumers
  • Test end-to-end freshness and sharing latency to meet agreed SLA targets
  • Register connector types and implement RBAC plus tenant-scoped authorization for all data-out paths
  • Implement metering hooks compatible with billing requirements for governed data-out flows

Requirements

  • Proven experience with 5+ years in data engineering using Python and cloud data platforms
  • Hands-on expertise with Google Cloud BigQuery for advanced analytics and data modeling
  • Solid background working with lakehouse table formats, including Apache Iceberg and Delta Lake
  • Demonstrated leadership skills to steer integration design, technical decisions, and delivery ownership
  • Strong track record executing projects that deliver data pipelines and adapters meeting SLAs
  • Advanced skills in Kafka/CDC patterns for real-time and near-real-time data movement
  • Deep understanding of data lake and lakehouse ingestion architecture on object storage
  • Excellent collaboration skills to align with platform, governance, and downstream consumer needs
  • Upper-Intermediate English proficiency (B2) for technical discussions and documentation

Nice to have

  • Apache Spark experience for validation and consumption testing
  • Knowledge of Databricks Unity Catalog for governed access patterns
  • Delta Lake expertise, including Delta Sharing setup and troubleshooting
  • Gen AI Assisted Development experience with tools such as Claude Code, GitHub Copilot, or Cursor
  • Snowflake Horizon Catalog experience for metadata governance and external table management

Похожие вакансии

Другие вакансии EPAM