Responsibilities:
Maintain our data platform with timely and quality data
Plan and execute data platform expansion to support the company’s growth and analytic needs. Define and extend our internal standards for style, maintenance, and best practices for a high-scale data platform
Build and maintain data pipelines from internal databases and SaaS applications
Write maintainable, performant code and create and maintain systems documentation
Desire to continually keep up with advancements in data engineering practices.
Solve technical problems of the highest scope and complexity
Implement the DataOps philosophy in everything you do
Design and develop code to extend the enterprise data model
Build trust in all interactions and with trusted data development
Collaborate with Analytics Engineers and Data Analysts to drive efficiencies for their work
Collaborate with other functions to ensure data needs are addressed
Requirements:
Python/Scala/Java
Data Processing(Apache Spark/Apache Flink/Apache Beam, etc.) and Pipeline orchestration (Apache Airflow, Apache NiFi end, etc)
Message Brokers and distributed streaming platform
DB (SQL/NoSQL/Open formats)
Dev Tools (Git/Docker/Kubernetes/Jira/etc)
Infrastructure/Cloud (AWS/GCP/Azure/etc).
Data Privacy/Security
Algorithms/Data Structure/Data Modeling
Data Development Principles/Process
Writing Requirements/Documentation
CI/CD/CD
Practical familiarity with DataOps, Data Lake, Data Mart, and DataMesh/Data Hub principles
Nice to have: AWS Redshift, AWS Glue, S3, EMR, DMS, QuickSight, MWAA, etc