4,305 open roles
Senior AI Engineer with Spark, AWS Services
Job description
We are seeking a Senior AI Engineer with Spark and AWS Services expertise to join the RBQM Production Pod within the program. You will build and maintain data pipelines that power AI/GenAI applications for Risk-Based Quality Management in clinical trials. This role focuses on RAG document ingestion, vector indexing, and building data APIs for AI applications.
Responsibilities
Design and build RAG document ingestion pipelines (chunking, embedding, vector indexing) for clinical trial quality data Build and manage vector databases (AWS Open. Search) for RAG-powered AI workflows Develop batch and streaming ETL/ELT pipelines from scratch for unstructured clinical data (PDF, DOCX, clinical reports) Build and expose data APIs for AI application consumption Optimize chunking strategies, embedding generation, and retrieval performance for RAG architectures Manage data quality, lineage, and governance for AI/ML data pipelines Deploy and maintain AWS data infrastructure (S3, Lambda, Glue, Athena, Step Functions, DynamoDB) Collaborate with Data Scientists and Backend Developers in an integrated pod team Requirements 5+ years of hands-on data engineering experience at scale Expertise in RAG document ingestion pipelines (chunking, embedding, vector indexing) Proficiency in AWS Open.
Search as a vector database for RAG workflows Advanced proficiency in Python, including SQL and Spark SQL Skills in unstructured data transformation (PDF, DOCX) for RAG/LLM applications Familiarity with AWS Services: S3, Lambda, Glue, Athena, Bedrock, Step Functions, API Gateway, Cloud. Watch, DynamoDB Knowledge of containerization with Docker Capability to build custom pipelines from scratch, beyond configuring out-of-the-box services Proficiency in English at a B2+ level Nice to have Background in pharmaceutical or life sciences domain Familiarity with Snowflake, Pinecone (vector DB alternative) Knowledge of Sage.
Maker processing jobs Skills in CI/CD tools (Jenkins, Git/Bitbucket) and infrastructure tools (CDK or Terraform) Understanding of clinical data standards (CDISC, ADaM, SDTM)
Description copied from EPAM Systems's careers page. Read the full posting before you apply.
More jobs at EPAM Systems
Senior Functional Tester
EPAM Systems· Remote (Brazil; Argentina; Chile; Colombia; Mexico)Accounting Manager
EPAM Systems· Sao Paulo, BrazilOMS Lead Engineer
EPAM Systems· Lisbon, Distrito de Lisboa, PortugalSenior DevOps Engineer - Azure
EPAM Systems· Remote (Armenia; Azerbaijan; Georgia; Kazakhstan; Kyrgyzstan; Uzbekistan)Lead Cloud Engineer
EPAM Systems· Remote (Armenia; Georgia; Kazakhstan; Kyrgyzstan; Uzbekistan)