4,194 open roles
Data & AI Reliability Engineering Consultant/Architect
Job description
We are hiring Data & AI Reliability Engineering Consultants and Architects to lead client-facing discovery, presales and advisory work on the reliability of data platforms and AI systems. The role is a consultant: the person owns the conversation with the client, shapes the solution and the proposal, and then guides the engineering team that builds it.
The consultant covers three connected areas: observability and SRE practice, data platform reliability (pipelines, quality, lineage, cost), and AI/LLM system reliability (evaluation, telemetry, guardrails).
Responsibilities
qualify the request, run client workshops, define scope, assumptions and estimates Run discovery and maturity assessments of a client's observability, data reliability and AI operations; deliver findings and a prioritised roadmap Write the solution part of proposals and RFP responses; present and defend it to client technical and business stakeholders Design target architectures for observability and reliability of data platforms and AI/LLM workloads, including tool selection and migration paths Define SLOs, SLIs and error budgets for data products and AI services; translate them into alerting, incident and governance processes Build the business case: cost of incidents, tooling cost optimisation, expected effect of the change Act as the trusted advisor for client engineering leads, SDMs and directors during the engagement Lead the first phase of delivery after a won deal, then hand over to the engineering team while staying accountable for the solution Review the work of engineers, set technical standards (alert-as-code, dashboards-as-code, IaC), unblock decisions Turn project experience into reusable assets: offerings, accelerators, assessment frameworks, reference architectures Mentor engineers growing towards consulting; take part in technical interviews Represent Data & AI externally: talks, articles, vendor partnerships Requirements 7+ years in engineering, of which 2+ years in a client-facing role: consultant, solution architect, presales engineer or technical lead with direct client ownership Proven presales record: led discovery or assessment workshops, produced estimates and proposals, presented to senior stakeholders Able to structure an ambiguous client problem into scope, options, trade-offs and a recommendation, in writing and live English B2+ with confident spoken delivery; can run a workshop and handle objections without support Hands-on background with at least one enterprise observability platform: New Relic, Datadog, Splunk, Dynatrace, Grafana stack or Elastic Open.
Telemetry, distributed tracing, metrics and log pipelines; alert design, event correlation and noise reduction SRE practice in production: SLO/SLI, error budgets, incident management, postmortems Cloud (AWS, Azure or GCP), Kubernetes, Terraform or other IaC; Python or similar for automation Understands how modern data platforms work and fail: Databricks, Snowflake or a cloud-native equivalent; orchestration (Airflow or similar); batch and streaming Data reliability practice: data quality checks, freshness and volume monitoring, lineage, pipeline SLAs, cost observability Working understanding of LLM application architecture (RAG, agents, model gateways) and what has to be measured: quality evaluation, latency, token cost, drift, guardrails Experience instrumenting or operating at least one AI/ML workload in production, or designing such a solution for a client Self-driven and reliable on commitments: owns deadlines for proposals and client deliverables without supervision Comfortable switching between several presales and one delivery engagement Uses AI assistants in daily engineering and documentation work Nice to have Vendor certifications: New Relic, Datadog, Splunk, Dynatrace; Databricks or Snowflake; cloud architect level (AWS, Azure, GCP) Data observability tooling: Monte Carlo, Soda, Great Expectations, Databricks Lakehouse Monitoring, Unity Catalog LLM observability and evaluation tooling: Langfuse, Lang.
Smith, Arize, MLflow, Open. Telemetry GenAI conventions AIOps and ITSM integration: Service. Now, Pager. Duty, event correlation engines Fin. Ops for observability and data platforms; licence and ingestion cost optimisation AI security and governance basics: guardrails, red teaming, data masking Domain experience in retail, finance or manufacturing Public profile: conference talks, articles, community leadership
Description copied from EPAM Systems's careers page. Read the full posting before you apply.
More jobs at EPAM Systems
Senior Full-Stack Engineer (Python+React)
EPAM Systems· Remote (Serbia)Manager/Senior Manager, Delivery Management - Application Security
EPAM Systems· London, England, UKSenior Data Scientist
EPAM Systems· Remote (Argentina; Brazil; Chile; Colombia; Mexico)Senior End-User Support Engineer
EPAM Systems· Warsaw, Masovian Voivodeship, PolandSenior Cloud & Infrastructure Engineer (Azure & Terraform)
EPAM Systems· Remote (Canada)