826 open roles
Senior Quantization Engineer - Edge AI Model Optimization
Job description
We at NXP have an environment that fosters innovation. Our team has technology experts who understand the big picture and mentors who coach passionate professionals to work on the most exciting challenges. We share responsibilities in everything we do, where every point of view is valued. Join us!
Job Summary
We are seeking a highly skilled Edge AI Engineer/Scientist with a strong theoretical foundation in AI and solid software engineering expertise to contribute to our Edge AI Model Optimization program. While the primary focus of this role is on model quantization, the scope also includes complementary optimization strategies such as speculative decoding, pruning, and other methods for ensuring highly efficient on-device deployment.
You will work at the forefront of innovation, bridging the gap between research and practice, focusing on CNNs, Large Language Model (LLM) and Vision Language Model (VLM) optimization to bring advanced capabilities to NXP’s Ara2 family of NPUs, directly supporting the future of on‑device intelligence. If you want to the future of efficient on‑device AI, this is the place to be.
Job Responsibilities
Actively survey the latest research (NeurIPS, ICLR, CVPR) on model optimization/compression, focusing particularly on neural network quantization, but also including other techniques like speculative decoding, pruning, etc. Prototyping: Develop and adapt state-of-the-art methods to NXP’s hardware constraints, building POCs to showcase the effectiveness of these techniques on NXP HW.
Production Implementation: Translate research prototypes into robust, optimized production code (C++/Python), ensuring strict memory and compute efficiency standards. Systems Integration: Document algorithmic tradeoffs, derive deployment recipes, and mentor the engineering team on numerical methods and optimization. Cross-Functional Leadership: Act as the technical bridge between AI Research, Hardware Engineering and other teams, providing quantified guidance on how choices impact model accuracy and performance.
IP Generation: Contribute to NXP’s intellectual property portfolio through patents and technical publications. Job Qualifications Required Background Education: MSc or Ph.D (is a plus) in Computer Science, Electrical Engineering, or Mathematics with a focus on Machine Learning or Deep Learning. AI Expertise: Proven practical experience in AI/ML with a deep understanding of CNN architectures and Generative AI (Transformers, LLMs, VLMs, etc.)
Technical Stack: Strong hands-on experience with Py. Torch, ONNX, and model conversion/optimization pipelines. Software Engineering: Proficient in Python and C++ and best development practices. Embedded Mindset: Familiarity with the constraints of embedded systems (latency, power, memory bandwidth) and how code interacts with underlying hardware.
Preferred Advanced AI: Experience with state-of-the-art quantization techniques for discriminative and generative AI (e.g., GPTQ, Spin. Quant, etc). Hardware Acceleration: Experience with NPUs, device-level profiling, and diagnosing memory bottlenecks. Kernel Development: Experience with custom kernel development is a plus.
Compilers: Knowledge of MLIR or TVM is a significant plus. More information about NXP in India... #LI-29f4
Description copied from NXP Semiconductors's careers page. Read the full posting before you apply.
More jobs at NXP Semiconductors
Intern (Engineering & IT)
NXP Semiconductors· Kuala LumpurSenior ASIC Design Engineer
NXP Semiconductors· PuneSr Manager - Mass Market Technical Leader - Americas
NXP Semiconductors· 2 LocationsBootROM Lead Architect
NXP Semiconductors· 2 LocationsSenior Validation Engineer (f/m/d)
NXP Semiconductors· Hamburg
More jobs in Hyderabad
Software Engineering Senior Advisor - HIH - Evernorth
Accredo by Evernorth· Hyderabad, IndiaSoftware Engineering Associate Advisor - HIH - Evernorth
Cigna· Hyderabad, India