We are looking for a Lead Data Software Engineer to build and optimize an Apache Iceberg-based data lakehouse on AWS while strengthening automation and platform tooling. You will partner with senior engineers to improve reliability and governance and help scale tenant onboarding—apply to make a measurable impact.
Responsibilities
- Build and optimize an Apache Iceberg lakehouse on AWS for performance and reliability
- Execute Iceberg optimization including partitioning, compaction, snapshot retention, orphan-file removal, and schema evolution
- Improve data quality and reliability across Bronze-layer landing zones
- Enhance configuration-management tooling used across the lakehouse platform
- Build self-service tenant-onboarding automation for landing zones, IAM, and monitoring
- Maintain the automated GitHub-to-S3 deployment pipeline
- Contribute to PR-based validation workflows and uphold code review standards
- Evaluate Iceberg-to-Snowflake integration options and document trade-offs
- Collaborate with senior engineers to align implementations with platform architecture
Requirements
- 5+ years of data engineering experience with cloud platforms
- Hands-on Amazon Web Services experience
- Advanced Apache Iceberg expertise including optimization and schema evolution
- Advanced data lakehouse architecture knowledge across landing zones and layers
- Strong leadership skills to guide delivery under senior oversight
- Proven project experience improving reliability and data quality in production systems
- Strong automation skills for configuration management and CI/CD deployment pipelines
- Strong collaboration skills for PR-based workflows and cross-engineering alignment
- Upper-Intermediate English proficiency (B2)
- Clear written communication skills for technical documentation and trade-off analysis
Nice to have
- Snowflake integration experience
- AWS Glue development experience