Job description
亮点 / Highlights: 🌍 Global Reach Open to global tech talent, not geo-restricted Explicitly requires experience in globally distributed teams across time zones Remote/hybrid options — very appealing to international candidates 🤖 AI + Infrastructure Dual Focus Rare intersection role: not pure AI research, not pure infra — the bridge between both Covers the hottest areas: LLM infrastructure (training, fine-tuning, RLHF, efficient serving) GPU cluster management & heterogeneous compute (A100/H100, CUDA, NCCL) — seriously hard-core
岗位职责 / Duties: Technical Leadership & Architecture • Define and drive the technical vision and long-term architecture strategy for AI/MLplatforms, model training/serving infrastructure, and large-scale data systems • Lead the design of next-generation AI infrastructure including model training pipelines,feature stores, real-time inference systems, and GPU/accelerator cluster management • Make high-impact design decisions on distributed systems that handle millions of QPSand petabyte-scale data processing • Establish and champion engineering best practices for AI system reliability, performanceoptimization, and cost efficiency AI/ML Systems & Infrastructure • Architect and build scalable ML platforms that accelerate model development, training,evaluation, and deployment lifecycles • Design high-performance inference serving systems with low-latency requirements forreal-time AI applications • Drive innovation in LLM (Large Language Model) infrastructure, including fine-tuning,RLHF pipelines, and efficient serving frameworks • Lead the development of foundational infrastructure components: container orchestration,service mesh, observability, and CI/CD for ML workloads • Optimize compute resource utilization across heterogeneous hardware (GPU, TPU,custom accelerators) Execution & Delivery • Hands-on contribution to the most complex and ambiguous technical challenges in AI andinfrastructure • Own end-to-end delivery of large-scale, cross-team projects from concept to production • Drive technical feasibility assessments for new AI product initiatives and infrastructureinvestments • Identify and resolve systemic technical risks, performance bottlenecks, and reliabilityconcerns Influence & Mentorship • Partner with senior leadership (VP/Director level) to align AI/infrastructure strategy withbusiness objectives • Drive cross-functional collaboration across AI research, engineering, product, andplatform teams • Mentor and coach senior engineers, fostering a culture of technical excellence andinnovation • Raise the engineering bar through design reviews, code reviews, and architecturalguidance • Influence the broader tech community through open-source contributions, publications, orconference talks