Job description
ACG_3671_JOB Our client is a leading AI infrastructure software company, seeking an experienced professional to join their team. Team Leadership & Delivery Lead, coach, and develop a team of 4–5 GPU/HPC engineers, promoting a culture of technical excellence, ownership, and continuous improvement. Plan and coordinate sprint activities, allocate workloads effectively, and ensure the timely delivery of high-quality GPU software components.
Conduct regular one-on-one meetings, provide constructive feedback, and support the professional growth and career development of team members. Convert high-level technical objectives from senior engineers or engineering management into clear, executable tasks for the team. Monitor team progress, proactively identify and escalate blockers, and provide clear status updates to relevant stakeholders.
Hands-on Technical Contribution Develop production-grade GPU kernel code using CUDA, HIP, or OpenCL for AI training and inference workloads. Lead code review activities and uphold coding standards, GPU optimization practices, and software quality across the team. Perform performance profiling and optimization of GPU kernels, including memory hierarchy usage and execution efficiency.
Contribute directly to technically complex features and resolve challenging issues that require deep expertise in GPU systems.
Requirements
Minimum Qualifications Bachelor’s degree in Computer Science, Computer Engineering, or a related discipline. Strong proficiency in C++ and Python. Hands-on experience with CUDA, HIP, or OpenCL. Experience optimizing GPU memory hierarchies, including shared memory, registers, memory coalescing, and occupancy. Familiarity with deep learning frameworks such as Py.
Torch or Tensor. Flow, including their interaction with GPU computing environments. Proven experience leading or mentoring a small technical engineering team of at least two members. Strong analytical and problem-solving capabilities, with the ability to diagnose and resolve complex GPU software issues. Good written and verbal communication skills to support team coordination, documentation, and stakeholder communication.
Preferred Qualifications
Master’s degree or Ph.D. in Computer Science, Computer Engineering, Artificial Intelligence, or a related field. At least two years of professional experience developing GPU system software. Experience with distributed GPU computing, multi-GPU coordination, or parallel runtime systems. Understanding of AI model architectures and their impact on GPU workload design, such as attention mechanisms and matrix operations.
Demonstrated track record of owning end-to-end delivery for a team, module, or technical workstream. Experience using profiling tools such as Nsight Compute, Nsight Systems, or AMD ROCm Profiler. Contributions to open-source GPU/HPC projects or publications at relevant conferences such as PPoPP, HPDC, SC, MICRO, or similar venues.
Contact: Thao Phan and Thu Giang Van Due to the immense number of applicants, only shortlisted candidates will be contacted