Job description
Responsibilities
Fine-tune and adapt open-source speech models on our proprietary call audio Build training and evaluation pipelines for multilingual and code-mixed speech Own model quality against metrics that matter operationally — entity and numeric accuracy, latency to first response — not just aggregate error rates Design and run data curation at scale: pseudo-labelling, speech enhancement, quality filtering on messy real-world audio Work with our linguist on text normalisation and pronunciation handling Evaluate candidate architectures, make the call with evidence, and ship the result to production with the platform team Requirements Must have skills: 3–4 years in ML, with at least 18 months on speech or audio specifically Strong Python and Py.
Torch; comfortable reading a paper and implementing it Hands-on experience fine-tuning at least one production speech model Solid grasp of speech fundamentals — mel-spectrograms, acoustic models and vocoders, encoder-decoder vs transducer architectures, evaluation methodology, sampling rates and what they cost you Understanding of how modern speech systems are actually built: self-supervised encoders, neural audio codecs, LM-based generation, flow matching Experience with genuinely messy audio, not only clean benchmark datasets Nice to have: Ne.
Mo, ESPnet, Speech. Brain, or Coqui Telephony-band or contact-centre audio Multilingual or code-switched speech work LoRA/PEFT, distributed training Open-source contributions or publications in speech
More jobs at Rezo.ai
Enterprise Sales Manager
Rezo.ai· Noida, Uttar Pradesh, IndiaQC Analyst - Bengali
Rezo.ai· Noida, Uttar Pradesh, IndiaCustomer Success Manager
Rezo.ai· Noida, Uttar Pradesh, IndiaSenior Backend Engineer - Speech Platform
Rezo.ai· Noida, Uttar Pradesh, IndiaTalent Acquisition Executive
Rezo.ai· Noida, Uttar Pradesh, India