Job description
Key Responsibilities
- Own end-to-end resolution of L3 incidents via Service. Now, including RCA and closure within SLA
- Perform deep technical analysis across Pega, Spring Boot services, integrations, and database layers
- Utilize: o Splunk for log analysis, correlation, and triage o Dynatrace/Data. Dog for performance analysis, and dependency mapping
- Support and monitor: o Batch jobs, schedulers, queue processors, listeners, and application health
- Raise, track, and manage defects in JIRA / JTFM, including: o Detailed RCA documentation o Mapping production issues to backlog items o Driving fixes to closure with Dev teams
- Participate in daily defect triage calls and ensure alignment between production issues and JTFM tracking
- Execute and validate standard, emergency, and release-related changes
- Collaborate with L2, Dev, Infra, and vendor teams for incident triage, escalation, and resolution
- Maintain runbooks, KT artifacts, and audit-compliant documentation
- Identify proactive monitoring gaps, alert tuning, and automation opportunities Mandatory Skills & Experience
- Strong Pega Platform experience (CSA/PCSA or equivalent) o Tracing, clipboard, rules debugging, job schedulers, queues
- Solid Java / Spring Boot troubleshooting (APIs, microservices)
- Hands-on experience with: o Dynatrace (RCA, dashboards, alerting) o Splunk (log queries, analysis, dashboards)
- Experience in L3 Production Support Model
- Strong working knowledge of: o Service. Now (Incident / Problem / Change) o JIRA / JTFM defect management lifecycle
- Experience in RCA-driven issue resolution and MTTR management