1,433 open roles
Research Engineer - [Seed Model - Infra - LLM/VLM Inference Optimization (Kernel & Compiler)]
Job description
The Seed Infrastructures team oversees the distributed training, reinforcement learning framework, high-performance inference, and heterogeneous hardware compilation technologies for AI foundation models.
Responsibilities
- Design, implement, and optimize high-performance GPU kernels for large-scale LLM/VLM inference workloads, including attention, GEMM, and other compute- and memory-intensive operators.
- Develop and tune inference kernels in CUDA and Triton, and drive end-to-end performance optimization of production inference systems at scale.
- Conduct in-depth performance analysis and profiling to identify bottlenecks across the inference stack, from kernel level to serving level.
- Collaborate with research and infrastructure teams to land kernel- and compiler-level optimizations in production inference systems.
Qualifications
Minimum Qualifications
- Bachelor's degree or above in Computer Science, Electrical Engineering, or a related field.
- Strong proficiency in C/C++ and Python; solid foundations in algorithms, data structures, and systems programming.
- Hands-on experience in LLM/VLM inference optimization with demonstrated impact on latency, throughput, or serving cost.
- Hands-on experience writing and optimizing GPU kernels in CUDA and/or Triton.
- Deep understanding of GPU architecture (memory hierarchy, occupancy, instruction throughput) with solid optimization experience.
Preferred Qualifications
- Experience with ML compiler internals (e.g., Triton, MLIR, LLVM).
- Contributions to related open-source projects (e.g., Triton, vLLM, SGLang, FlashAttention, CUTLASS).
- Publications in relevant venues (e.g., MLSys, OSDI, ASPLOS).
Description copied from ByteDance's careers page. Read the full posting before you apply.
More jobs at ByteDance
Talent Acquisition Partner (Global AI To B) - BytePlus
ByteDance· San Jose, California, United StatesCorporate Finance Associate - Group FP&A
ByteDance· Hong Kong (China), Hong Kong Island, Hong Kong, ChinaTalent Acquisition Partner (Global AI To B) - BytePlus
ByteDance· London, England, United KingdomDatacenter Operation Engineer (DCO) - Infrastructure Engineering
ByteDance· Kulai, Johor, MalaysiaBackend Software Engineer (SRE) - Cloud Infrastructure
ByteDance· Singapore
More jobs in San Jose
Software Architect - PNS AI Governance
TikTok· San Jose, California, United StatesJob Posting Title Senior Data Management Project Engineer
Hitachi· San Jose, California, United StatesSenior Software Development Engineer
Expedia Group· USA - California - San Jose· $199k – $279kSr Technical Architect
Johnson Controls International· San Jose-San Jose-Costa RicaStaff Cell Design Engineer
Lyten· San Jose, CA· $141k – $212k