We are seeking a Lead Data Science Engineer to design and deliver reusable data sharing adapters and governed access patterns across a cloud lakehouse and external data platforms. You will guide high-impact architecture decisions, ensure reliable data and AI workflows, and help the team ship secure integrations.
Responsibilities
- Design a reusable lakehouse write layer with dual-format metadata to support multiple consumers
- Build and validate ingestion pipeline patterns from object storage to an analytical warehouse for structured operational data
- Implement change-data-capture patterns for real-time and near-real-time data movement into the lakehouse
- Develop dependency-aware bookkeeping and data lineage tracking patterns across pipelines
- Ensure adapter code is modular, version-controlled, tested, and reusable across new integrations
- Configure external table definitions for shared lakehouse data products in Snowflake catalogs
- Validate zero-copy read access from Snowflake to shared Iceberg and Delta tables without data movement
- Implement tenant-scoped access controls aligned with external catalog governance requirements
- Implement and certify a Delta Sharing adapter for live data sharing to Databricks consumers
- Configure Delta Sharing endpoints and manage sharing agreements for multiple tenants
- Validate consumer access via supported clients while meeting freshness and latency expectations
- Register connector types in a governed connector registry and enforce auditable RBAC on data-out paths
- Implement metering hooks compatible with governed billing requirements for external data flows
Requirements
- 5+ years of data science or ML engineering experience with production Python and pandas
- Experience building RAG applications using embeddings and retrieval pipelines
- Experience writing SQL for analytical data workflows and validation
- Strong technical leadership skills to drive architecture decisions and mentor peers
- Proven project delivery skills across multi-system data integration workstreams
- Solid software engineering skills in modular design, testing, debugging, and Git workflows
- Hands-on cloud platform skills with GCP, AWS, or Azure fundamentals
- Strong LLM fundamentals knowledge including tokenization, attention, context windows, and sampling
- Practical prompt engineering skills with structured outputs and few-shot techniques
- Robust evaluation and monitoring skills including drift, performance, and hallucination detection
- Strong communication and collaboration skills across engineering and data stakeholders
- Upper-Intermediate English proficiency (B2, Upper-Intermediate)
- Active experience using AI-assisted development tools such as Claude Code, GitHub Copilot, or Cursor
Nice to have
- Google Cloud Platform experience with BigQuery and object storage patterns
- Large Language Models (LLM) API integration experience including streaming, rate limits, and cost controls
- Vector database experience with indexing, chunking strategies, and retrieval tuning
- Experience optimizing LLM latency and cost using caching, batching, and model routing