Lead Data Science Engineer

EPAM·Argentina, Colombia, Mexico, Brazil, Chile·Удалённо·вчера

We are seeking a Lead Data Science Engineer to design and deliver reusable data sharing adapters and governed access patterns across a cloud lakehouse and external data platforms. You will guide high-impact architecture decisions, ensure reliable data and AI workflows, and help the team ship secure integrations.

Responsibilities

  • Design a reusable lakehouse write layer with dual-format metadata to support multiple consumers
  • Build and validate ingestion pipeline patterns from object storage to an analytical warehouse for structured operational data
  • Implement change-data-capture patterns for real-time and near-real-time data movement into the lakehouse
  • Develop dependency-aware bookkeeping and data lineage tracking patterns across pipelines
  • Ensure adapter code is modular, version-controlled, tested, and reusable across new integrations
  • Configure external table definitions for shared lakehouse data products in Snowflake catalogs
  • Validate zero-copy read access from Snowflake to shared Iceberg and Delta tables without data movement
  • Implement tenant-scoped access controls aligned with external catalog governance requirements
  • Implement and certify a Delta Sharing adapter for live data sharing to Databricks consumers
  • Configure Delta Sharing endpoints and manage sharing agreements for multiple tenants
  • Validate consumer access via supported clients while meeting freshness and latency expectations
  • Register connector types in a governed connector registry and enforce auditable RBAC on data-out paths
  • Implement metering hooks compatible with governed billing requirements for external data flows

Requirements

  • 5+ years of data science or ML engineering experience with production Python and pandas
  • Experience building RAG applications using embeddings and retrieval pipelines
  • Experience writing SQL for analytical data workflows and validation
  • Strong technical leadership skills to drive architecture decisions and mentor peers
  • Proven project delivery skills across multi-system data integration workstreams
  • Solid software engineering skills in modular design, testing, debugging, and Git workflows
  • Hands-on cloud platform skills with GCP, AWS, or Azure fundamentals
  • Strong LLM fundamentals knowledge including tokenization, attention, context windows, and sampling
  • Practical prompt engineering skills with structured outputs and few-shot techniques
  • Robust evaluation and monitoring skills including drift, performance, and hallucination detection
  • Strong communication and collaboration skills across engineering and data stakeholders
  • Upper-Intermediate English proficiency (B2, Upper-Intermediate)
  • Active experience using AI-assisted development tools such as Claude Code, GitHub Copilot, or Cursor

Nice to have

  • Google Cloud Platform experience with BigQuery and object storage patterns
  • Large Language Models (LLM) API integration experience including streaming, rate limits, and cost controls
  • Vector database experience with indexing, chunking strategies, and retrieval tuning
  • Experience optimizing LLM latency and cost using caching, batching, and model routing

Похожие вакансии

Другие вакансии EPAM