← Live feed

TTS Model development

Budget: $300.0 FIXED / ⭐ 0.00 (0) India

convolutional-neural-network, artificial-neural-networks, deep-neural-networks, python, deep-learning, artificial-intelligence

Preferred qualifications

  • Experience: Expert
## Multilingual Speech AI / TTS Engineer We are looking for an experienced **Speech AI / TTS Engineer** to develop a high-quality multilingual **Text-to-Speech (TTS)** system for 9 Indian languages: 1. Telugu 2. Hindi 3. Tamil 4. Marathi 5. Bengali 6. Gujarati 7. Kannada 8. Malayalam 9. Indian English The project involves **end-to-end TTS model development**, including open-source dataset research, data preparation and cleaning, integration of studio-recorded data that we provide, model architecture research, experimentation, training, evaluation, and development of high-quality voices for each supported language. The engineer will be expected to evaluate suitable modern TTS architectures and clearly document why particular approaches were selected or rejected. The goal is to develop a robust, scalable multilingual TTS system capable of producing natural, high-quality speech across all 9 languages. ### Key Responsibilities * Research and evaluate suitable open-source datasets for the 9 target languages. * Review dataset licensing and ensure appropriate usage. * Clean, filter, normalize, and prepare speech and text data for training. * Integrate and preprocess studio-recorded speech data provided by us. * Develop and train multilingual and/or multi-speaker TTS models. * Research and evaluate modern TTS architectures such as **VITS, FastSpeech, diffusion/flow-based TTS, and other state-of-the-art approaches**. * Conduct systematic model experiments and hyperparameter optimization. * Develop efficient training and inference pipelines using PyTorch. * Optimize GPU utilization and training efficiency on the infrastructure provided by us. * Evaluate speech quality, pronunciation, intelligibility, naturalness, and consistency across languages. * Perform detailed failure analysis and iterate on the models. * Deliver reproducible training, evaluation, and inference pipelines. * Maintain comprehensive technical documentation throughout the project. ### Documentation Requirements Comprehensive documentation is a major requirement of this project. Documentation should cover: * Dataset sources and licenses * Data collection and preprocessing * Audio quality filtering * Text normalization and phonemization * Language-specific preprocessing * Model architecture decisions * Experiments and results * Training configurations * Evaluation methodology and metrics * Failure analysis * Model comparison * Final model selection and rationale * Reproducibility instructions * Training and inference procedures ### Infrastructure We will provide the required **GPU infrastructure** for model training. The engineer will be responsible for efficiently configuring and utilizing the infrastructure, optimizing training workloads, managing experiments, and delivering reproducible training and inference pipelines. ### Required Experience The ideal candidate should have strong practical experience with: * **PyTorch** * TTS and speech synthesis models * Multilingual and/or multi-speaker TTS * Audio preprocessing and speech datasets * Speaker representations and embeddings * GPU-based model training * Distributed or large-scale training * Model evaluation and benchmarking * Modern TTS architectures such as: * VITS * FastSpeech / FastSpeech 2 * Diffusion-based TTS * Flow-based TTS * Transformer-based TTS * Other modern state-of-the-art TTS architectures Experience with **Indian languages**, particularly Telugu, Hindi, Tamil, Marathi, Bengali, Gujarati, Kannada, Malayalam, or Indian English, is highly desirable. ### Please Include in Your Proposal 1. Examples of TTS models you have personally **trained or developed**. 2. Details of any **multilingual or multi-speaker TTS** systems you have worked on. 3. Experience working with **Indian languages or low-resource languages**, if applicable. 4. Which TTS architecture you would investigate first for this project and **why**. 5. Your approach to multilingual training and handling language-specific pronunciation and text normalization. 6. Your approach to dataset preparation, quality filtering, and evaluation. 7. Examples of previous work involving large-scale GPU training. 8. Links to relevant GitHub repositories, papers, demos, or deployed systems, where available. We are looking for an engineer who can take ownership of the **complete TTS development pipeline**, from dataset preparation and architecture selection through training, evaluation, optimization, and final model delivery.
Open job

AI proposal draft

Generate a short cover letter to copy into the offer. Says you are interested and ready to work.

Sign in to generate an AI proposal draft.

Log in