← Zákazky

Build Lightweight Indian Speech-to-Text (STT) Models (less than 100 MB) for Multiple Languages

Rozpočet: $200.0 FIXED / ⭐ 0.00 (0) India

artificial-neural-networks, pytorch, machine-learning, python

Project Overview We are looking for an experienced AI/ML Engineer with expertise in Automatic Speech Recognition (ASR), Speech-to-Text (STT), multilingual speech processing, and deep learning to build a suite of production-ready lightweight STT models focused on Indian languages. The objective is to develop **individual language-specific Speech-to-Text models**, with **each model being under 100 MB**, while maintaining high transcription accuracy, low latency, and efficient CPU inference. The deliverables should include the complete training pipeline, fine-tuning pipeline, documentation, and inference code so that our team can continue improving the models after project completion. --- Project Requirements 1. Lightweight Models Each language-specific model should: * Be under **100 MB** in size. * Be optimized for CPU inference. * Support ONNX export (preferred). * Be deployable on Linux. * Have low memory usage and low inference latency. * Be suitable for future Android deployment. Please explain any trade-offs if the target model size requires sacrificing recognition accuracy. --- 2. Languages Develop individual Speech-to-Text models for the following Indian languages: * Hindi * Telugu * Tamil * Kannada * Malayalam * Bengali * Marathi * Gujarati * Punjabi * Assamese Additional Indian languages are welcome if supported by available datasets. Each language should have its own independent model, rather than a single multilingual model. --- 3. Speech Recognition Quality The models should provide: * High transcription accuracy. * Low Word Error Rate (WER). * Accurate recognition of proper nouns and common vocabulary. * Good robustness across different speakers, genders, accents, and speaking speeds. * Stable transcription quality in typical real-world environments. Reasonable trade-offs between model size and accuracy are acceptable, provided they are clearly justified. --- 4. Code-Mixed Speech Support The models should be capable of recognizing common English words that naturally occur in Indian language conversations. Examples include: ```text "Meeting is at five PM." "Payment complete అయింది." "Project submit कर दिया." "Please call me वापस." "Invoice send చేయండి." ``` The models should accurately transcribe both the native language and frequently used English words where appropriate. Please describe your proposed approach for handling code-mixed speech. --- 5. Open-Source Datasets Use publicly available datasets wherever possible. Examples include: * AI4Bharat datasets * OpenSLR * Mozilla Common Voice * Google FLEURS * IndicSUPERB * Hugging Face speech datasets * Other permissively licensed datasets While the preference is to build the models using open-source datasets, **commercial speech datasets may also be recommended and purchased if they provide significant improvements in transcription accuracy, speaker diversity, accent coverage, or language coverage.** If recommending commercial datasets, please clearly provide: * Dataset name * Licensing information * Estimated pricing * Technical justification * Expected improvement over open-source datasets --- 6. Training Pipeline Provide complete reproducible training code including: * Dataset download scripts * Audio preprocessing * Audio normalization * Voice activity detection (if required) * Text normalization * Tokenizer generation * Training scripts * Validation pipeline * Checkpoint saving * Resume training * Evaluation metrics (WER, CER, etc.) The project should be reproducible from scratch. --- 7. Fine-Tuning Pipeline One of the primary deliverables is a complete fine-tuning pipeline. It should allow us to: * Improve existing language models. * Continue training from checkpoints. * Fine-tune using smaller datasets. * Adapt models for domain-specific vocabulary. * Add additional speakers and accents where applicable. * Generate updated inference models after fine-tuning. Please provide clear documentation for the complete fine-tuning workflow. --- 8. Documentation Provide comprehensive documentation covering: * Environment setup * Dependency installation * Dataset preparation * Training * Fine-tuning * Evaluation * ONNX export (if applicable) * Offline inference * Streaming inference (if supported) * Model architecture overview * Hyperparameters * Expected hardware requirements * Troubleshooting guide The documentation should enable another engineer to reproduce the complete pipeline without additional guidance. --- 9. Deliverables The final deliverables should include: * Complete source code * Training scripts * Fine-tuning scripts * Dataset download scripts * Dataset preprocessing scripts * Configuration files * Trained model weights * Exported inference models * ONNX models (preferred) * Example inference applications * Documentation * License information for all third-party assets --- Preferred Skills * Automatic Speech Recognition (ASR) * Speech-to-Text (STT) * Deep Learning * PyTorch * ONNX * Hugging Face * NVIDIA NeMo * Whisper * Parakeet * Conformer * Citrinet * wav2vec 2.0 * MMS * Audio Signal Processing * Multilingual NLP * Indian Language Processing --- Proposal Requirements Please include the following in your proposal: 1. Similar Speech-to-Text or ASR projects you have completed. 2. Languages you have previously worked with. 3. Model architecture you recommend and why. 4. Expected model size for each language. 5. Expected Word Error Rate (WER) or Character Error Rate (CER). 6. Whether streaming inference is supported. 7. Hardware required for training. 8. Estimated timeline. 9. Estimated budget. 10. Public GitHub repositories, research papers, demos, or previous work. --- Important Notes * Preference should be given to open-source models, datasets, and libraries wherever possible. * Commercial speech datasets may be proposed if they provide meaningful improvements in transcription accuracy or language coverage. Any recommendations should include licensing details, estimated costs, and a clear technical justification. * Each language should be delivered as an **independent model under 100 MB**. * The final solution should be production-ready and structured for long-term maintenance. * Preference will be given to candidates who have demonstrable experience building lightweight, high-accuracy Speech-to-Text systems for Indian languages and who can justify architectural decisions with respect to model size, latency, deployment efficiency, and transcription accuracy.
Otvoriť na Upwork