← Jobb

Python Developer Needed to Build Automated VTT Caption & Transcript Alignment Pipeline

Budget: $500.0 FIXED / ⭐ 5.00 (2) USA

api-integration, restful-api, python, ffmpeg, automatic-speech-recognition, natural-language-processing, machine-learning, automation, data-processing, json, video-processing, subtitling, timestamps

Preferred qualifications

  • Experience: Intermediate
We are looking for an experienced **Python developer to build an automated pipeline that generates accurately timed WebVTT (.vtt) captions from professionally edited transcripts and existing video files.** We have approximately **100 videos totaling about 120 hours**, each with a corresponding edited transcript in DOCX or PDF format. The Goal We already have the correct caption text. **What we need is the timing.** The system should use a speech-to-text service such as **ElevenLabs Scribe** to generate word-level timestamps from each video's audio. The automatically transcribed words are not the final captions. Instead, the system should align those timestamped words with our professionally edited transcript and transfer the timing onto the edited text. The result should be: **Our exact edited transcript wording + accurate word timing = finished VTT captions.** What You Will Build The pipeline should: * Pair approximately 100 videos with their corresponding DOCX/PDF transcripts using a manifest. * Use **ffmpeg** to extract transcription-ready audio from each video. * Send audio to the **ElevenLabs Speech-to-Text API** and retrieve word-level timestamps and speaker diarization. * Support keyterms for proper nouns, names, organizations, and specialized terminology. * Extract and clean transcript text from DOCX and PDF files. * Remove page numbers, headers/footers, timestamps, line numbers, and other non-spoken material. * Normalize both the automatic transcription and edited transcript for comparison while preserving the original edited wording for display. * Align the two word streams even when filler words, false starts, grammar, names, numbers, or terminology differ. * Transfer timestamps from matching speech-to-text words to the edited transcript. * Interpolate timing for edited words that do not have a direct speech-to-text match. * Generate properly formatted caption cues. * Export standards-compliant **WebVTT (.vtt)** files. * Produce automated QC reports identifying files that need manual review. Alignment Is the Key Technical Challenge The edited transcripts intentionally differ from what was literally spoken. Editors have removed filler words and false starts and corrected grammar, names, and terminology. Therefore, this is **not a simple transcription project**. The developer needs to implement sequence/word alignment that tolerates insertions, deletions, and replacements without causing timing to drift. A Python `SequenceMatcher` approach may be used as a starting point, but we are open to a more robust alignment method. **Please explain your proposed alignment approach when applying.** Caption Requirements Generated captions should generally follow these rules: * Maximum 2 lines per cue * Maximum 42 characters per line * Approximately 84 characters maximum per cue * Minimum duration of approximately 1 second * Maximum duration of approximately 7 seconds * Maximum reading speed of approximately 21 characters per second * Prefer sentence and natural phrase boundaries * Break cues at significant pauses where appropriate * No overlapping captions * Maintain a small gap between consecutive cues where possible * Never rewrite or shorten the edited transcript simply to make a caption fit The **final VTT text must match our edited transcript exactly** after appropriate normalization. Quality Control The pipeline should automatically check: * VTT validity * Transcript-to-caption text completeness * Alignment/match percentage * Cue duration * Reading speed * Line length * Number of lines per cue * Overlapping cues * Incomplete speech-to-text results Alignment scores should be reported for every file and sorted from worst to best so questionable files can be reviewed first. The workflow should also make it easy to identify areas requiring manual review, including overlapping speakers or unusually poor alignment. Deliverables We need the **complete working system**, not simply the finished caption files. Final deliverables should include: * All Python source code * ffmpeg/audio extraction workflow * ElevenLabs API integration * DOCX/PDF transcript extraction and cleaning * Transcript-to-word-timestamp alignment system * Caption cue-generation logic * VTT generation * Automated QC/validation * Batch processing * Error handling and retries * Ability to rerun an individual video by ID * Requirements/dependency file * Installation instructions * Written runbook/documentation The system must be reusable so that future videos can be added without rebuilding the pipeline or rehiring the developer. Ideal Experience We are particularly interested in developers with experience in: **Python, ffmpeg, ElevenLabs or other speech-to-text APIs, NLP/text alignment, forced alignment, word-level timestamps, subtitle/caption generation, WebVTT/SRT, DOCX/PDF parsing, and batch-processing pipelines.** Previous work involving **forced alignment or matching edited transcripts against spoken audio** is especially relevant. When Applying Please tell us: 1. Your experience with similar captioning, speech-to-text, NLP, or forced-alignment projects. 2. How you would align an edited transcript with imperfect word-level speech-to-text output. 3. How you would prevent timing drift across long videos. 4. How you would identify and handle low-confidence alignments. 5. Your estimated timeline. 6. Your fixed-price estimate or estimated hours. If you have built a similar transcription/alignment pipeline, please include an example. **This is a software development project, not a manual transcription or captioning job.** The project is complete when we can add a new video and its edited transcript, run the pipeline, and automatically receive an accurately timed VTT file containing our exact edited wording, along with a QC report indicating whether the file passed or needs manual review.
Öppna på Upwork

AI proposal draft

Generate a short cover letter for this job. Edit before sending.

Sign in to generate an AI proposal draft.

Logga in