We are seeking a Senior Data Science Engineer to design and deliver reusable data-sharing adapters and governed access patterns across a cloud analytics lakehouse. You will build reliable integrations, apply strong engineering practices, and help teams consume data safely and efficiently.
Responsibilities
- Design a dual-format lakehouse write layer that supports Delta and Iceberg metadata over one physical dataset
- Build and validate ingestion patterns from object storage to an analytical warehouse for structured operational data
- Implement change-data-capture patterns with Kafka for real-time and near-real-time data movement
- Develop dependency-aware bookkeeping and data lineage tracking patterns across pipelines
- Ensure adapter code is modular, version-controlled, and reusable for new data source integrations
- Configure external table definitions to enable governed access for downstream platforms
- Validate zero-copy read access to Iceberg and Delta tables without unnecessary data movement
- Implement and test tenant-scoped access controls aligned with metadata governance needs
- Deliver a Delta Sharing adapter to provide live, zero-copy sharing to external consumers
- Configure sharing endpoint registration and manage sharing agreements for data products
- Test end-to-end data freshness and sharing latency against agreed SLA targets
- Register connector types in a connector registry and enable controlled data-out connectivity
- Implement RBAC and tenant-scoped authorization for all external data-out paths with auditability
- Add metering hooks compatible with billing for governed data-out flows
- Document integration patterns and operating procedures for reuse and support
Requirements
- 3+ years of data science or ML engineering experience using Python, pandas, and scikit-learn
- Experience building RAG applications, including embeddings and retrieval strategies
- Strong leadership skills to drive technical decisions and mentor peers across workstreams
- Proven project ownership skills delivering reusable adapters and integration patterns end to end
- Advanced software engineering skills in modular design, testing, debugging, and Git workflows
- Strong SQL skills for analytical querying, validation, and pipeline support
- Solid cloud fundamentals across AWS, GCP, or Azure, including scalability and latency tradeoffs
- Hands-on monitoring skills for drift and performance, plus experiment tracking practices
- Strong AI tooling skills with AI-assisted development tools such as Claude Code, Cursor, or GitHub Copilot
- Upper-Intermediate English proficiency (B2) for clear technical communication with global stakeholders
Nice to have
- Google Cloud Platform experience with BigQuery and GCS
- Large Language Models (LLM) fundamentals, including tokenization and context windows
- Vector database experience for embeddings storage and retrieval
- LangChain or similar agent framework experience for tool use and orchestration
- LLM API integration experience, including streaming, rate limits, and cost management