Senior Video-Pipeline Engineer - Build an AI Short-Video Editing Service
Budget: -
HOURLY / FULL_TIME
⭐ 3.74 (5)
Israel
captions, audio-editing, video-editing, visual-effects, vfx-cleanup, motion-retouching, dialog-editing, subtitling
Preferred qualifications
- Experience: Expert
Title: Senior Video-Pipeline Engineer — Build an AI Short-Video Editing Service (Auto Captions, Cuts, Zooms, Music) — Multilingual incl. Hebrew & Arabic (RTL)
About us
We run a growing SaaS (web + mobile app) that creates and publishes social-media content for business clients. Today our video editing runs on a third-party API (Submagic/OpusClip-class tooling). We want to bring this capability in-house as our own service, integrated behind our existing backend.
What we're building
A server-side pipeline: it receives a video, edits it automatically, and returns a finished 1080×1920 MP4 via webhook. Capabilities, in priority order:
Milestone 1 — Captions core (the critical one):
Speech-to-text with word-level timestamps in English, Hebrew, French, Spanish, Portuguese, and Arabic. Hebrew is our primary acceptance bar (fast conversational speech), and note that two of the six languages — Hebrew and Arabic — are right-to-left.
Animated, styled captions rendered onto the video — template-based (colors, fonts, word-by-word highlight animation), with correct RTL rendering and bidirectional text handling (e.g., Latin words or numbers inside an RTL sentence). We provide the design specs of ~10 caption styles.
Silence/pause removal with 3 pacing levels, and optional removal of repeated takes.
Milestone 2 — Polish:
Auto punch-in zooms on key moments.
Background-music mixing: track by ID from our library, volume 1–100, fade in/out.
Milestone 3 — B-rolls:
AI-selected stock-footage overlays covering a configurable % of the video.
Milestone 4 — Long-form to clips:
Split a long video into N best short clips (min/max length params), with face-centered reframing to 9:16.
Integration requirements (non-negotiable)
REST API: POST /projects accepting ( video, language, captionTemplateId, magicZooms, brolls + %, removeBadTakes, silencePace, music: trackId, volume, fade , webhookUrl ) → async processing → webhook callback with a downloadable MP4 URL. We will share our current API contract so the service is a drop-in swap for our backend.
Two ingestion paths, both required: (a) direct upload from our apps — mobile included — so design for large files over flaky connections (resumable/chunked upload, e.g. tus or S3-compatible multipart); (b) pull from a URL we provide (our existing storage). Typical input: 1–10 min vertical videos, files can be large (2GB+).
Target turnaround: minutes, not hours.
Deployable to our cloud account (we own the infra and the code). Include a processing queue and basic monitoring.
You should have
Proven work with FFmpeg and programmatic video rendering (Remotion, Motion Canvas, or raw FFmpeg filter graphs).
Hands-on experience with ASR (Whisper large-v3, Deepgram, or similar) including word-level alignment; RTL experience (Hebrew/Arabic) is a big plus.
Backend + GPU-pipeline experience (Node or Python), API design, webhooks, queues, resumable uploads.
Portfolio links to similar video-automation work — required.
How we'll evaluate
Milestone-based contract. Milestone 1 acceptance: a blind side-by-side against our current provider on 10 real Hebrew videos — caption accuracy, timing, and visual quality judged by us — plus spot checks in the other five languages. We pay per accepted milestone.
Open job
AI proposal draft
Generate a short cover letter to copy into the offer. Says you are interested and ready to work.
Sign in to generate an AI proposal draft.
Log in