Job description
LM Studio is used by millions of people around the world to run AI on their own computers, and now with Bionic - also in the cloud. Our values prioritize putting the human in the center, and creating tools that we want to use ourselves, and recommend to our friends and family. As a team, we work with high technical intensity and personal responsibility.
We are looking for curious, self-motivated, creative, and technically excellent teammates to join us and build the future of human-AI interactions in software. The Role We are looking for an Inference Runtime Software Engineer to push forward LM Studio's inference stack on-device and in the cloud. You will integrate new inference engines and runtime capabilities, bring up new open-weight models and modalities, and optimize model execution for a wide range of CPU and GPU targets.
You will also contribute improvements to the open-source projects we build on.
Qualifications
- Significant experience building production ML systems, inference runtimes, or performance-sensitive infrastructure
- Strong programming ability in Python and C++
- Deep understanding of transformer architectures and the mechanics of model inference
- Experience profiling CPU or GPU workloads and reasoning about compute, memory, synchronization, and data movement
- Experience with Py. Torch and inference systems such as llama.cpp, MLX, Execu. Torch, vLLM, SGLang, or TensorRT-LLM
- Strong debugging instincts across model code, runtime internals, operating systems, and CPU or GPU execution
- Takes personal responsibility for the correctness and performance of their work Bonus Qualifications
- Past contributions to open-source inference runtime projects such as llama.cpp, MLX, Execu. Torch, vLLM, SGLang, or TensorRT-LLM Responsibilities
- Maintain and push forward our inference stack on-device and in the cloud
- Bring up new model architectures and multimodal models
- Improve latency, throughput, memory use, and reliability across CPU, CUDA, Metal, Vulkan, and ROCm runtimes
- Build runtime capabilities for model loading, batching, scheduling, caching, and distributed execution
- Benchmark and diagnose correctness and performance problems across the inference stack
- Contribute upstream to open-source projects such as llama.cpp and MLX Benefits
- Competitive salary and equity grants
- Great medical, vision, dental healthcare plans
- Catered team lunch / expensed dinners in the office
- Flexible PTO
- Flexible WFH
- Sun-drenched office in So. Ho in NYC
More jobs at Lm Studio
Marketing
Lm Studio· New York City· $150k – $250kFull Stack Software Engineer
Lm Studio· New York City· $150k – $350kSoftware Engineer, Agent Harness
Lm Studio· New York City· $150k – $350kGTM, Enterprise
Lm Studio· New York City· $170k – $270kSoftware Engineer, Application
Lm Studio· New York City· $175k – $275k
More jobs in New York City
V.I.E. - 18 months - XVA F/M - New York
Groupe BPCE· New York, États-Unis, InternationalMega-Mansion Estate Manager (200 staff +) with USA exp., UHNW Principal (Doha, Qatar)
educated-households· New York, USARegistered Dietitian
The Vernon Staffing Group· New York, United StatesSpoken English Coach (American Accent)
LeadInTop· New York, United StatesSaaS Lead Solutions Engineer
redtech-recruit· New York, United States· $150k – $200k