We are looking for a Senior Data Engineer to join our team. The Data Engineer is the most data-intensive engineering role on the engagement. If any pipeline drops data, positions drift, or signals compute incorrectly, calibration and UAT break. Correctness and operational robustness here are a prerequisite for everything else.
Responsibilities
- Execute all database migration sets spanning the Unified Data Store to ensure schema consistency across the platform
- Build and maintain the full suite of Source Adapters responsible for connecting external systems into the data platform
- Implement the Field Mapper, establishing per-project and per-franchise field bindings that normalize source system fields into the unified schema during ingestion
- Implement the Link Resolver, responsible for resolving CL-to-ticket, ticket-to-ticket, and case-to-defect/requirement links across all source families, feeding directly into the Attribution Resolver
- Implement the Attribution Resolver, tracing the change-to-ticket-to-area-to-test chain, alongside the Counted Signals Aggregator, which builds file-to-area and area-to-area counted association tables with path normalization and counting verification
- Build the Area Vocabulary and accompanying Translation Tables to support consistent area classification across the system
- Develop all Signal Catalogue computation jobs covering the eight core signals — area fragility, recency, recent failures, change-touch, coupling, windowed area change volume, testing alignment, validation recency, and defect impact/volume/age — ensuring provenance capture throughout
- Take ownership of data dictionary authoring across all storage components, documenting incrementally as new features are delivered
Requirements
- 3+ years of hands-on relevant experience in Python software engineering
- Background in AI Data Engineering, applying data engineering practices to support AI/ML-driven systems
- Practical experience with Apache Airflow for orchestrating and scheduling data workflows
- Working knowledge of Machine Learning concepts and their application within data systems
- Proven experience in data pipeline development, from design through implementation
- Hands-on experience with PostgreSQL for data storage and querying
- Experience designing idempotent ingest processes, including natural keys, upsert auditing, duplicate detection, and replay safety
- Experience integrating multiple heterogeneous data sources into a unified system
- Skilled in data quality and telemetry practices, including fill-rate counters, volume metrics, and reconciliation reporting
- Familiarity with Spec Driven Development methodology
- Decent communication skills with working English fluency (B2 level or higher) to understand business requirements and translate them into agentic architectures
Nice to have
- Experience developing API clients for Perforce or Code Hub
- Experience integrating with the JaaS (Jira) REST API
- Familiarity with test management APIs such as QMetry, Zephyr, or similar tools
- Experience building Snowflake connectors, including key-pair authentication and warehouse extract queries
- Knowledge of temporal data systems and as-of read patterns