Parth Lad
Building the future of voice AI
Data Scientist specializing in speech technologies, with a focus on building reliable, production-grade systems from the ground up. My work spans the full pipeline — from raw audio and video collection to deployment-ready speech solutions and I take pride in the quality and consistency of every layer in between.
Expertise
Key achievements
Projects
End-to-end fine-tuning of XTTS V2 on a curated Hindi speech corpus — covering data collection, preprocessing, speaker normalisation, and GPU training on RunPod infrastructure. Produces natural-sounding Hindi synthesis with accurate phoneme rendering, built to handle the phonetic complexity of Indic language inputs.
Real-time voice AI agent handling 14 Indian financial domains (banking, loans, insurance, investments, tax) across 11 languages including Hindi, Tamil, and Telugu. Features gender-adaptive personas, word-level transcription, and post-call stereo recording with turn-diarized CSV transcripts for training data collection.
A rule-based NLP engine that converts English text to Devanagari script at 86.6% accuracy, using a 7-tier lookup cascade from a curated 7,300+ word dictionary to a neural phoneme model. Handles edge cases like silent letters, suffix morphology, and multi-word phrases. Deployed as a Flask REST API with a live web interface.
Tech stack
Get in touch
I'm open to roles, collaborations, and consulting in speech AI, audio data engineering, and production ML systems. If you're working on something in voice tech — especially for Indic languages — let's talk.