About the Position
We are looking for a Senior Data Engineer with hands-on experience in Databricks, Snowflake, and either AWS or Azure to join our team. In this role, you will own the full lifecycle of our data warehouse, from ETL and ELT pipeline design and dimensional modeling to performance tuning and infrastructure optimization. You will work with modern data platforms to transform raw data into reliable, scalable, and actionable insights that support business decision making.
Responsibilities
- Develop, operate, optimize, test, and maintain the data warehouse, including ETL/ELT process development, cube development, database/performance administration, and dimensional table design
- Drive the full life-cycle of back-end development for the data warehouse
- Identify, design, and implement internal process improvements - redesigning infrastructure for scalability, optimizing data delivery, and automating manual processes
- Define data retention policies
- Build analytical tools that leverage the data pipeline to deliver actionable insight into key business metrics (operational efficiency, customer acquisition, etc.)
- Select and integrate tools for monitoring, managing, alerting on, and improving database performance
- Develop and implement automated processes to ensure uninterrupted database updates and correction of vulnerabilities
- Assemble large, complex datasets that meet functional and non-functional business requirements
Requirements
- 5+ years of experience or 5+ completed projects
- Advanced SQL and query optimization, with proficiency across popular database variations
- Python for data engineering and automation
- Cloud data platforms: Snowflake and/or Databricks
- Cloud services: AWS and/or Azure
- Data transformation and modeling with DBT
- ETL/ELT pipeline design, development, and maintenance
- Apache Spark (PySpark preferred)
- Workflow orchestration using Apache Airflow, Dagster, or another widely adopted orchestrator
- Relational database design and performance tuning (PostgreSQL, MySQL, SQL Server, Oracle, etc.)
- Data warehousing concepts and dimensional modeling
- Data management fundamentals: data modeling, data quality, metadata management, data warehouse/lake patterns, distributed systems
- Version control using Git and CI/CD practices
- Data governance, data quality, lineage, and observability practices
- Security and access control implementation in cloud data platforms
Nice to Have
- Apache Kafka, Apache Flink, Apache Beam
- Terraform, Kubernetes, Docker
- Apache Iceberg, Delta Lake, or Apache Hudi
- Real-time and event-driven architectures
- AWS Glue, Amazon MWAA, Azure Data Factory
- Data Mesh and Data Product concepts
- Machine learning data pipelines, feature stores, or AI development experience
- Streaming analytics and Change Data Capture (CDC) solutions (e.g., Debezium)