1,433 open roles
Senior Production Engineer - Applied Machine Learning
Job description
The mission of our AML team is to push next-generation recommendation-based algorithms and platform for the company. We also drive substantial impact for core businesses of the company. Currently we are looking for Production Engineers to join our team to support and advance that mission
Responsibilities
- System Stability & Production Management: Responsible for the production management and stability assurance of AML (Applied Machine Learning) training, inference, and storage systems. This covers core pipelines including scheduling and orchestration, K8s/GPU clusters, distributed training, online inference serving, and ParameterServer/NoSQL storage.
- Reliability Engineering: Build and maintain mechanisms for SLO/SLA, observability, alerting, On-call processes, fault diagnosis, auto-healing, disaster recovery, and incident reviews (post-mortems).
- Engineering Excellence: Drive engineering capabilities such as CI/CD, canary releases, auto-rollback, automated inspections, pre-flight checks, capacity forecasting, and elastic auto-scaling.
- Resource & Cost Management: Oversee resource governance across GPU/CPU/storage/network, including quota management, cost attribution, and performance tuning, to improve system availability, resource utilization, and overall R&D efficiency.
Qualifications
- Bachelor’s degree or above in Computer Science, Software Engineering, Artificial Intelligence, or related fields.
- Familiar with Linux and proficient in at least one of the following programming/scripting languages: Shell, Python, Go, or C++.
- Understanding of machine learning training/inference architectures, Kubernetes, GPU clusters, or distributed storage systems.
- Proven experience in online troubleshooting, performance analysis, and building automation platforms.
- Strong sense of responsibility, clear logical thinking, and the ability to drive the resolution of complex issues across cross-functional teams.
Preferred Qualification(s):
- Experience with large-scale training/inference/storage platforms, SLO governance, FinOps, NoSQL, or open-source infrastructure is highly preferred.
Description copied from ByteDance's careers page. Read the full posting before you apply.
More jobs at ByteDance
Talent Acquisition Partner (Global AI To B) - BytePlus
ByteDance· San Jose, California, United StatesTalent Acquisition Partner (Global AI To B) - BytePlus
ByteDance· London, England, United KingdomCorporate Finance Associate - Group FP&A
ByteDance· Hong Kong (China), Hong Kong Island, Hong Kong, ChinaBackend Software Engineer (SRE) - Cloud Infrastructure
ByteDance· SingaporeCustomer Success Manager - Lark Japan
ByteDance· Tokyo, Japan
More jobs in San Jose
Advanced AI for Industry & Society Staff TA - College of Engineering - Integrated Innovation Institute
Carnegie Mellon University· Silicon Valley, CAWarehouse Part Time Overnight
Lowe's· San Jose, CA (S San Jose) 1756· $40k – $41kRetail Manager - CLUB Membership
Bass Pro Shops· San Jose, CA· $70k – $83kSoftware Architect - PNS AI Governance
TikTok· San Jose, California, United StatesLead Organizer
Californians for Justice· San Jose, CA 95133· $82k