Rezo.ai logo

ML Engineer - Speech

Rezo.ai

Noida, Uttar Pradesh, IndiaFull-timePosted Sep 28, 2026

Job description

Responsibilities

Fine-tune and adapt open-source speech models on our proprietary call audio Build training and evaluation pipelines for multilingual and code-mixed speech Own model quality against metrics that matter operationally — entity and numeric accuracy, latency to first response — not just aggregate error rates Design and run data curation at scale: pseudo-labelling, speech enhancement, quality filtering on messy real-world audio Work with our linguist on text normalisation and pronunciation handling Evaluate candidate architectures, make the call with evidence, and ship the result to production with the platform team Requirements Must have skills: 3–4 years in ML, with at least 18 months on speech or audio specifically Strong Python and Py.

Torch; comfortable reading a paper and implementing it Hands-on experience fine-tuning at least one production speech model Solid grasp of speech fundamentals — mel-spectrograms, acoustic models and vocoders, encoder-decoder vs transducer architectures, evaluation methodology, sampling rates and what they cost you Understanding of how modern speech systems are actually built: self-supervised encoders, neural audio codecs, LM-based generation, flow matching Experience with genuinely messy audio, not only clean benchmark datasets Nice to have: Ne.

Mo, ESPnet, Speech. Brain, or Coqui Telephony-band or contact-centre audio Multilingual or code-switched speech work LoRA/PEFT, distributed training Open-source contributions or publications in speech

More jobs at Rezo.ai

See all openings at Rezo.ai