Data & AI Reliability Engineering Consultant/Architect

EPAM·Ukraine·Удалённо·сегодня

We are hiring Data & AI Reliability Engineering Consultants and Architects to lead client-facing discovery, presales and advisory work on the reliability of data platforms and AI systems. The role is a consultant: the person owns the conversation with the client, shapes the solution and the proposal, and then guides the engineering team that builds it.

The consultant covers three connected areas: observability and SRE practice, data platform reliability (pipelines, quality, lineage, cost), and AI/LLM system reliability (evaluation, telemetry, guardrails).

Responsibilities

  • Lead the technical side of presales: qualify the request, run client workshops, define scope, assumptions and estimates
  • Run discovery and maturity assessments of a client's observability, data reliability and AI operations; deliver findings and a prioritised roadmap
  • Write the solution part of proposals and RFP responses; present and defend it to client technical and business stakeholders
  • Design target architectures for observability and reliability of data platforms and AI/LLM workloads, including tool selection and migration paths
  • Define SLOs, SLIs and error budgets for data products and AI services; translate them into alerting, incident and governance processes
  • Build the business case: cost of incidents, tooling cost optimisation, expected effect of the change
  • Act as the trusted advisor for client engineering leads, SDMs and directors during the engagement
  • Lead the first phase of delivery after a won deal, then hand over to the engineering team while staying accountable for the solution
  • Review the work of engineers, set technical standards (alert-as-code, dashboards-as-code, IaC), unblock decisions
  • Turn project experience into reusable assets: offerings, accelerators, assessment frameworks, reference architectures
  • Mentor engineers growing towards consulting; take part in technical interviews
  • Represent Data & AI externally: talks, articles, vendor partnerships

Requirements

  • 7+ years in engineering, of which 2+ years in a client-facing role: consultant, solution architect, presales engineer or technical lead with direct client ownership
  • Proven presales record: led discovery or assessment workshops, produced estimates and proposals, presented to senior stakeholders
  • Able to structure an ambiguous client problem into scope, options, trade-offs and a recommendation, in writing and live
  • English B2+ with confident spoken delivery; can run a workshop and handle objections without support
  • Hands-on background with at least one enterprise observability platform: New Relic, Datadog, Splunk, Dynatrace, Grafana stack or Elastic
  • OpenTelemetry, distributed tracing, metrics and log pipelines; alert design, event correlation and noise reduction
  • SRE practice in production: SLO/SLI, error budgets, incident management, postmortems
  • Cloud (AWS, Azure or GCP), Kubernetes, Terraform or other IaC; Python or similar for automation
  • Understands how modern data platforms work and fail: Databricks, Snowflake or a cloud-native equivalent; orchestration (Airflow or similar); batch and streaming
  • Data reliability practice: data quality checks, freshness and volume monitoring, lineage, pipeline SLAs, cost observability
  • Working understanding of LLM application architecture (RAG, agents, model gateways) and what has to be measured: quality evaluation, latency, token cost, drift, guardrails
  • Experience instrumenting or operating at least one AI/ML workload in production, or designing such a solution for a client
  • Self-driven and reliable on commitments: owns deadlines for proposals and client deliverables without supervision
  • Comfortable switching between several presales and one delivery engagement
  • Uses AI assistants in daily engineering and documentation work

Nice to have

  • Vendor certifications: New Relic, Datadog, Splunk, Dynatrace; Databricks or Snowflake; cloud architect level (AWS, Azure, GCP)
  • Data observability tooling: Monte Carlo, Soda, Great Expectations, Databricks Lakehouse Monitoring, Unity Catalog
  • LLM observability and evaluation tooling: Langfuse, LangSmith, Arize, MLflow, OpenTelemetry GenAI conventions
  • AIOps and ITSM integration: ServiceNow, PagerDuty, event correlation engines
  • FinOps for observability and data platforms; licence and ingestion cost optimisation
  • AI security and governance basics: guardrails, red teaming, data masking
  • Domain experience in retail, finance or manufacturing
  • Public profile: conference talks, articles, community leadership

Похожие вакансии

Другие вакансии EPAM