Speech AI · Data Science · Production Systems

Parth Lad

Building the future of voice AI

Data Scientist specializing in speech technologies, with a focus on building reliable, production-grade systems from the ground up. My work spans the full pipeline — from raw audio and video collection to deployment-ready speech solutions and I take pride in the quality and consistency of every layer in between.

80%Testing effort reduced
Pipeline throughput gain
1yr+Experience

Expertise

🎙️
Speech Recognition & TTS

End-to-end ASR pipelines with WER benchmarking across six platforms including AssemblyAI, Sarvam, Azure, and Meta Omnilingual. Custom XTTS V2 fine-tuning and OmniVoice voice cloning deployed for Indic languages and enterprise BFSI use cases.

XTTS V2OmniVoiceVoice cloningWER analysisAssemblyAI
🔊
Audio Data Engineering

Production pipelines covering segmentation, format conversion, denoising, speaker diarization, and DNSMOS-based quality scoring at scale. Handles both single-speaker TTS corpora and two-person conversational audio including podcasts and phone call recordings.

FFmpegPydubLibrosaDiarizationConversational audio
🚗
Wake Word Detection

Automotive-grade wake word systems for enterprise clients covering the full pipeline from data collection through signal processing to production deployment. Optimised for low-latency, always-on detection in resource-constrained edge environments.

AutomotiveEdge deploymentLow-latencySignal processing
🏗️
Dataset Workflows

Studio-grade recording setups with HITL annotation — automated pipelines handle first-pass processing while human reviewers validate uncertain outputs before data enters training. Covers single-speaker TTS corpora, conversational datasets, and automated quality frameworks.

Single-speaker corporaConversational datasetsHITL workflowsQuality frameworks
👁️
Computer Vision

Object detection and tracking with YOLO-based models for real-time and batch inference workloads. Includes frame extraction pipelines, annotation workflow management, and dataset curation for consistent model training.

YOLOObject detectionFrame extractionAnnotation pipelines
☁️
Infrastructure & DevOps

Linux-native development with Docker containerisation, Azure GPU VMs for model training, and RunPod for speech processing pipelines including OmniVoice voice cloning. AWS and Cloudflare R2 for scalable audio storage across production workloads.

Azure VMRunPodDockerAWSCloudflare R2

Key achievements

01
Custom XTTS V2 model trained for an Indic language — native speaker-led data collection, preprocessing, and fine-tuning for production-grade Gujarati TTS.
TTS
02
Speech pipeline optimisation — overhauled FFmpeg/Pydub-based segmentation and format conversion pipelines, achieving a 3× throughput improvement on production workloads.
3×↑

Projects

Hindi TTS — XTTS V2
Built

End-to-end fine-tuning of XTTS V2 on a curated Hindi speech corpus — covering data collection, preprocessing, speaker normalisation, and GPU training on RunPod infrastructure. Produces natural-sounding Hindi synthesis with accurate phoneme rendering, built to handle the phonetic complexity of Indic language inputs.

XTTS V2Coqui TTSRunPodPythonSpeaker normalisation
BFSI Voice Agent
Built

Real-time voice AI agent handling 14 Indian financial domains (banking, loans, insurance, investments, tax) across 11 languages including Hindi, Tamil, and Telugu. Features gender-adaptive personas, word-level transcription, and post-call stereo recording with turn-diarized CSV transcripts for training data collection.

LiveKit AgentsTurn DetectionNoise CancellationAudio post-processingRAG
English → Devanagari Transliteration Engine
Built

A rule-based NLP engine that converts English text to Devanagari script at 86.6% accuracy, using a 7-tier lookup cascade from a curated 7,300+ word dictionary to a neural phoneme model. Handles edge cases like silent letters, suffix morphology, and multi-word phrases. Deployed as a Flask REST API with a live web interface.

PythonNLPFlaskREST APINeural phoneme model

Tech stack

Speech & Audio
XTTS V2 / Coqui TTSAssemblyAI, Sarvam, ElevenLabsAzure Speech, Google STTMeta Omnilingual ASRFFmpeg, Pydub, LibrosaDNSMOS
ML & Data Science
Python, Pandas, NumPyWhisperYOLO object detectionLLM IntegrationRAG pipelines & vector memoryinference & cost optimizationSegFormerFastAPIGradio / Streamlit
Infrastructure
Linux / DebianDocker, VMsAzureRunPodAWS, Cloudflare R2Git, GitHub
Datasets & Benchmarking
WER / MER / CER analysisDNSMOSRTFx & latency profilingTurn-diarized transcripts
Web & Backend
React, Node.js, PythonMongoDB, PostgreSQLREST APIs

Get in touch

I'm open to roles, collaborations, and consulting in speech AI, audio data engineering, and production ML systems. If you're working on something in voice tech — especially for Indic languages — let's talk.