Developer Wanted — Automated Book-Reel Video Generator (JSON → 135 Videos)
Бюджет: $500.0
FIXED /
⭐ 5.00 (3)
Cyprus
api-integration, python, javascript, automation, ffmpeg, video-processing
What I Need
I market books with short vertical videos (Reels/TikTok style, 9:16, 1080×1920). Today the content is fully prepared by an AI pipeline — for each book I receive a JSON file with 135 ready-made scenes in 3 fixed video formats. I want a tool that turns that JSON into the 135 finished videos automatically.
I need a developer to build a simple hosted web app that does the following:
Input: I upload the scenes JSON file + the book cover image.
Image generation: The tool generates one still image per scene via an image-generation API, using the ready-made image prompt included in the JSON for every scene, with the book cover as a style reference so all images share one consistent look.
Review step (required): Before any video is rendered, the tool shows me ALL generated images in a grid (grouped by format). For each image I can click Approve or Regenerate (regenerate = re-run the same prompt, optionally with a tweak field). Rendering only starts when every image is approved.
Rendering: The tool renders all 135 videos from 3 fixed templates (specs below). No creative decisions — the templates are fixed, the JSON supplies the variables.
Output: 135 MP4s named F1-Scene01.mp4 … F3-Scene45.mp4, downloadable as a ZIP or synced to a cloud folder (GDrive).
The same tool will be reused for many books and multiple languages(English, German, French, Italian), so it must be book-agnostic: new JSON + new cover in, videos out.
The 3 Video Templates
Format 1 — Dialogue Slideshow (45 videos per book)
8 slides, exactly 3 seconds each (~24 sec total), rendered as video.
One background image for the whole video. The image must show two characters/entities as two separate visual anchors (left/right or top/bottom) — e.g., a man and a spaceship facing each other.
Slides 1–7: dialogue from the JSON. Every slide displays one exchange pair — both lines at once: line a overlaid on Speaker A's side/zone of the image, line b on Speaker B's side/zone (the JSON delivers each slide as an a/b pair). Readable white text with soft shadow/outline (see examples).
Slide 8: always the book cover image, full frame.
Text placement: the two text zones can be fixed template positions. To guarantee the characters actually sit in those zones, the image step for Format 1 should either (a) enforce composition via the prompt + my review step, or (b) generate/segment the two characters separately and composite them into fixed positions. Propose your approach in your application.
Format 2 — Scrolling Text Video (45 videos per book)
60–90 seconds. Still image on the top ~40% of the frame; below it a black panel with white text scrolling slowly and continuously upward (text disappears behind the image edge, new lines enter from the bottom).
Scroll speed auto-calibrated from word count so the total duration lands between 60 and 90 seconds at a comfortable reading pace.
Text comes from the JSON (text field). The scroll is the only animated element.
Format 3 — Book Page Scroll Video (45 videos per book)
60–90 seconds. Full-frame colored background image; centered white "book page" panel with the scene text scrolling slowly (only animated element); scroll speed auto-calibrated like Format 2.
Top: the hook line from the JSON in a colored rounded text box.
In the text: the JSON delivers the content as ordered segments — narration (plain) and speech (with speaker name + a light highlight color, e.g., "light blue", "light pink", "light green"). Speech segments are rendered with that soft highlight color behind the text, like a marker.
Bottom left: eye icon + view-count number from the JSON. Bottom: small book icon + book title from the JSON (no author).
Example videos/images of all three formats are attached, plus a sample JSON showing the exact schema.
Technical Notes
Image generation must run through an API (e.g., Flux via Replicate/fal.ai, or comparable). You tell me which services you need — I will open the accounts myself and give you access for the build. Style-reference support from the cover image is required.
Rendering: your choice of stack (e.g., Remotion, FFmpeg pipeline, or a rendering API like Creatomate/Shotstack) — propose what you'd use and why.
Hosting: simple and cheap (this is an internal tool for one user). Password-protected page is enough.
Fonts/styling of the templates will be matched once to the attached examples during setup, then stay fixed.
The tool must fully support content in English, German, French, and Italian (the scene JSON may arrive in any of these four languages; text rendering, fonts, and special characters like ä/ö/ü/ß/é/à/ç must display correctly in all templates).
Deliverables
The working hosted tool (input → image review → rendering → output)
A short usage guide (1 page is fine)
Handover of the code/repo
A realistic estimate of the running costs per 135 videos (image API + rendering + hosting)
In Your Application, Please Include
Similar things you've built (template video rendering, AI-image pipelines, batch automation)
Your proposed stack for image generation + rendering, and your approach to the Format 1 two-character/text-placement problem
Fixed price (or price range) and timeline for a working v1
Ongoing: whether you're available for small improvements after launch
To be clear about the goal: after handover, I run new books through the tool myself — new JSON + new cover in, 135 videos out, in any of the four supported languages, with no developer needed. Possible follow-up work is limited to maintenance and small improvements.
Відкрити на Upwork