Member of Technical Staff - Inference
Job description
Sail builds the world's most efficient software for inference (processing LLM tokens) and agent hosting (cloud VMs). Together, our technologies allow our customers to deploy AI agents at extreme scale for their most ambitious work. In this role, you'll own token processing down to the lowest layers of the stack. You'll do things like: develop a new request scheduling strategy, achieve better communication/computation overlap, investigate novel schemes for increasing cache hit rates, or identify a better way to benchmark inference performance.
What you’ll do
- Modify and extend state-of-the-art inference engines like vLLM and SGLang, and work on our own internal engine.
- Understand every microsecond of GPU time spent during a forward pass. You'll be able to explain every kernel launch on an nsys profile.
- Design and implement exotic parallelism schemes to work with "interesting" hardware topologies.
- Write and debug GPU kernels to excel in specific regimes, such as cascade attention https://flashinfer.ai/2024/02/02/cascade-inference.html What we’re looking for
- Strong understanding of core LLM mechanics, like KV cache, mixture-of-experts, prefill vs. decode phases.
- Interest in MLSys research - great ideas like speculative decoding and sparse attention come from research, that we need to follow closely.
- Familiarity with modern, tile-based GPU programming, e.g. Triton, CUTLASS, ThunderKittens, etc. Or an interest in learning these!
- Great interpersonal and technical communication. Please don't use LLMs to write prose. We desk-reject slopful cover letters and resumes. Interview process
- Meet the CTO, who will ask about your experience, and share as much technical detail about Sail as you want to hear. This is the first step because we respect your time.
- Share an online whiteboard with a team member and work through a technical problem. We spend a lot of time at whiteboards, building intuition about complex systems together. It's a great way for us to see how you communicate technically, and a even better way for you to see what working at Sail is like.
- Come in to Sail's SF office for an interview day. Meet the whole team, and work on a bunch of problems that closely simulates the work we do daily. We'll also ask you to give us a 20-30min 'chalk talk' about an interesting problem you've worked on before.
- Offer. Once the team decides we want to work with you, we make a strong offer quickly and will be quite persistent over email/text/calls :) LIFE AT SAIL We work out of a beautiful, sunny office in downtown San Francisco. All meals are on us (and actually great; SF is a food paradise!). Everyone gets a Studio Display (or two) at their desk. We are serious about investing in anything that saves us time or energy. There are six different ways to make coffee or tea in the office. A friendly (hypoallergenic) black cat named Coco visits occasionally.
More jobs at Sail Research
Member of Technical Staff - Monetization
Sail Research· San FranciscoMember of Technical Staff - Sailboxes
Sail Research· San FranciscoStrategic Finance Lead
Sail Research· San Francisco· $200k – $300kMember of Technical Staff - Agent Engineering
Sail Research· San Francisco· $200k – $300kMember of Technical Staff - Distributed Systems
Sail Research· San Francisco· $200k – $300k
More jobs in San Francisco
Founding AI Engineer in SF
eth· San Francisco, United StatesSenior Frontend UX Engineer
Hubs· San Francisco - Hybrid / Global - Remote, United StatesChief of Staff
Talentry· San Francisco, United StatesSimulation Scenario Specialist
Your IT Recruiter· San Francisco, United StatesSan Francisco, CA: Court Supervisor
ePlay Digital· San Francisco, United States