EPAM Systems logo
EPAM Systems

4,192 open roles

Senior ML/Evaluation Engineer

Remote (Czech RepublicSlovakia)Full-timePosted Sep 29, 2026

Job description

We are building an Enterprise Agent Development Platform, a production-grade, cloud-native ecosystem that enables engineering teams to define, orchestrate, deploy, and observe AI agents at scale. The platform standardizes agent development across the organization. The initiative spans agent framework design, runtime architecture, marketplace integration, CI/CD automation, and enterprise-grade observability, reducing agent development from months to days while enforcing consistent security, quality, and governance standards.

Responsibilities

AWS Agent. Core Evaluation (on-demand mode for CI/CD gates, online mode for production sampling) LLM-as-judge evaluator design (built-in Agent. Core evaluators — helpfulness, correctness) Custom code-based Lambda evaluators (Python — deterministic checks) Evaluation levels (TRACE for per-response, TOOL_CALL for per-invocation, SESSION for workflow) OTel spans from AWS Agent.

Core Observability as evaluation input Enterprise evaluation standard authoring (mandatory dimensions, pass/fail criteria) Requirements 5+ years ML engineering or AI platform engineering LLM evaluation framework design and implementation Custom evaluator implementation for deterministic quality checks CI/CD deployment gate design for ML model or agent quality Nice to have AWS Agent.

Core Evaluation API hands-on (Create. Evaluation, Get. Evaluation. Result) AWS Bedrock Guardrails for PII detection evaluator integration Cloud. Watch metrics output from Agent. Core Evaluation for online mode

Description copied from EPAM Systems's careers page. Read the full posting before you apply.

More jobs at EPAM Systems

See all openings at EPAM Systems