319 open roles
Site Reliability Engineer (SRE) - Ads / Monetization Platform
Job description
Role Summary
As a Site Reliability Engineer (SRE), you will build and operate highly available, globally distributed advertising/monetization services. You will improve reliability, scalability, and operability through automation, observability, incident management, and sound engineering practices.
Key Responsibilities
design reviews, capacity planning, launch, deployment, operations, and continuous improvement. Build and operate highly available services across multiple regions/data centers; improve resilience, latency, and scalability. Develop automation and tooling to reduce toil (deployment, remediation, runbooks, self-healing) using scripting and software engineering best practices.
Define and implement SLOs/SLIs/SLAs; create dashboards and alerting to track service health (availability, latency, errors, saturation). Lead sustainable incident response: triage, mitigation, root-cause analysis (RCA), and blameless postmortems with actionable follow-ups. Collaborate with software engineering, security, and compliance stakeholders to meet data governance and regulatory requirements.
Requirements
Must-have Qualifications 3+ years of experience in SRE, DevOps, systems engineering, or production operations for large-scale services. Strong coding skills in one language: Python or Go or C++ (Java acceptable). Solid Linux/Unix fundamentals: processes, memory/CPU, filesystems, permissions, and troubleshooting. Networking fundamentals in cloud environments: TCP/IP, DNS, HTTP/HTTPS, load balancing, basic security concepts.
SQL proficiency and experience with data workflows/ETL is a plus for ads/analytics-related systems. Strong communication, ownership mindset, and ability to work effectively across global teams.
Preferred Qualifications
Experience supporting advertising, recommendation, or high-traffic consumer internet platforms. Hands-on experience with cloud platforms (AWS/GCP/Azure) and infrastructure-as-code (Terraform/Ansible). Experience with containers and orchestration (Docker, Kubernetes). Observability experience with tools such as Prometheus, Grafana, ELK/Splunk, Open.
Telemetry. Experience operating large data systems (streaming, distributed storage/compute) and performance tuning.
Description copied from Two95 International Inc.'s careers page. Read the full posting before you apply.
More jobs at Two95 International Inc.
Informatica mdm developer | Remote
Two95 International Inc.· IN (Remote)Helpdesk Technician - IT Vulnerability - 10 years experience
Two95 International Inc.· Newtown Square, PA, USSMP/E & z/OSMF Experts | Remote(EST/CST) - Long Term Contract
Two95 International Inc.· US (Remote)Z/os Experts | Remote (EST/CST) - Long Term Contract
Two95 International Inc.· US (Remote)z/OS Network Systems Administrator | Remote(EST/CST) - Long Term COntract
Two95 International Inc.· US (Remote)
More jobs in Kuala Lumpur
Sales Associate
Tapestry· Kuala Lumpur, MYS (My Kuala Lumpur Klcc - Coach)Chinese Merchant Support
Xendit· Kuala Lumpur, Malaysia; Taipei, TaiwanManager, Account Management
AIG· Kuala LumpurTechnical Specialist, Emerging Technologies
PlanGrid· APAC - Malaysia - Kuala Lumpur - Kuala Lumpur, Avenue 10Senior Associate - Digital Audit
PwC Ireland· Kuala Lumpur