4,194 open roles
Senior AI Reliability Engineer
Job description
EPAM's Operational Intelligence practice is developing a new capability called AI Reliability Engineering (AIRE) , which applies SRE principles and cloud-native practices throughout the lifecycle of production AI/ML and LLM systems. As clients transition their GenAI applications and agentic systems from pilot to production, they find that traditional APM provides no visibility into token latency, cost per request, semantic drift, or hallucinations.
This role exists to close that gap. Your work will involve instrumenting, monitoring, and hardening production AI systems, establishing AI-native service level objectives, and creating accelerators and reference implementations that the practice can reuse across accounts. Note that this is an engineering position rather than an L1/L2 support role, and it does not require a 24/7 on-call rotation.
What You'll Get A brand-new discipline within EPAM where you shape the approach rather than follow an existing playbook Focus on engineering work without 24/7 on-call responsibilities Sponsored certification and training programs (Anthropic/Claude, Databricks, AI & Data Observability learning paths) Exposure to multiple clients and a direct route into presales and solution engineering Responsibilities Add AI telemetry to production LLM, RAG, and agentic applications using Open.
Telemetry and APM-native AI monitoring tools Deploy distributed tracing across multi-model chains, agent workflows, and retrieval-augmented generation pipelines to identify systemic latency and failure points Establish and track AI-native SLIs and SLOs, including time to first token (TTFT), throughput, error and refusal rates, cost per request, semantic drift, hallucination boundaries, and contextual accuracy Implement structured semantic logging and prompt/response monitoring to support quality analysis Create and maintain evaluation loops for output quality and safety using golden sets, LLM-as-a-judge methods, and Ragas/Deep.
Eval-style frameworks, integrating them into CI/CD and runtime environments Monitor token-based cloud spend, model API rate limits, and quota usage while driving AI cost optimization efforts Set up AI gateways to manage API load balancing, failover, and fallback models across multiple LLM providers Build guardrails covering prompt injection and jailbreak filtering, output compliance, and bias and safety constraints Develop detection, triage, restoration, and problem management workflows for AI incidents, incorporating autonomous AI agents into root cause analysis to parse logs, generate hypotheses, and correlate state changes Enable rollback, canary, and fail-safe patterns for model, prompt, and configuration releases, while maintaining reproducibility through versioning of data, code, prompts, and models Develop practice accelerators, reference architectures, and internal training materials, and contribute to presales activities and client assessments Requirements 4+ years of experience in SRE, DevOps, platform, or observability engineering, with hands-on exposure to production AI/ML or LLM workloads Strong grasp of SRE fundamentals, including Golden Signals, SLI/SLO definition, error budgets and burn rate, incident lifecycle, and ITIL basics Strong Python skills for building instrumentation, automation, and evaluation tools Hands-on experience with Open.
Telemetry and at least one APM/observability platform such as New Relic, Datadog, Grafana LGTM stack, Splunk, or Elastic Production experience with at least one cloud platform (Azure preferred, AWS or GCP also acceptable) and Kubernetes Solid understanding of LLM application architecture, including prompts, embeddings and vector stores, RAG, and agent orchestration tools like Lang.
Chain/Lang. Graph or similar Awareness of MLOps concepts, including model lifecycle (training vs. inference), model endpoints, containerization, and deployment/rollback patterns Experience with Infrastructure as Code using Terraform, along with CI/CD tools such as Azure DevOps, GitLab CI, or GitHub Actions B2+ English proficiency, since the role involves direct client interaction and requires clear technical communication in both writing and speech Nice to have Familiarity with AI-specific observability and evaluation tools such as Traceloop/OpenLLMetry, Langfuse, Arize Phoenix, Ragas, Deep.
Eval, or MLflow Experience with distributed inference serving at scale, including vLLM, KServe, Ray Serve, Kubernetes-native LLM orchestration, and GPU capacity planning Knowledge of AI security practices, including the OWASP LLM Top 10, prompt injection defense, and guardrail frameworks like Ne. Mo Guardrails or Llama Guard Experience with Databricks (including Mosaic AI/MLflow) or Azure AI Foundry Relevant certifications such as Anthropic Claude, Azure AI Engineer, AWS ML Specialty, or Databricks GenAI Background in Fin.
Ops for AI workloads, particularly token and GPU cost modeling Experience in Data Reliability Engineering, covering data quality and pipeline SLOs, since AI reliability depends on data reliability Prior mentoring or team lead experience
Description copied from EPAM Systems's careers page. Read the full posting before you apply.
More jobs at EPAM Systems
Senior Full-Stack Engineer (Python+React)
EPAM Systems· Remote (Serbia)Manager/Senior Manager, Delivery Management - Application Security
EPAM Systems· London, England, UKSenior Data Scientist
EPAM Systems· Remote (Argentina; Brazil; Chile; Colombia; Mexico)Senior End-User Support Engineer
EPAM Systems· Warsaw, Masovian Voivodeship, PolandSenior Cloud & Infrastructure Engineer (Azure & Terraform)
EPAM Systems· Remote (Canada)