Presentation website for a TV channel with AI-generated news and voiceover (ElevenLabs API)
Budget: $300.0
FIXED /
⭐ 5.00 (14)
Kazakhstan
react-js, python, api-integration
Gewenste kwalificaties
- Ervaring: Expert
===================================================
1. PROJECT OVERVIEW
===================================================
Goal: build a presentation (demo) website for a TV channel that showcases automated news text generation and voiceover based on a short news thesis.
Usage scenario:
1. The user enters a short news thesis (e.g., "Ronaldo has arrived in Shymkent").
2. The site uses the Gemini API to generate a coherent news text, sized for an on-air voiceover of approximately 20-30 seconds.
3. The user selects a voice and voiceover style.
4. The site uses the ElevenLabs API to generate an audio recording of the text.
5. The user listens to the result in the browser (and optionally downloads the file).
Project conditions: the site is a demo, no user accounts or authentication. Text generation and voiceover must support three languages - Russian, Kazakh, English. Export only in mp3 format. No budget limit on API requests. Hosting - client's own node on Jelastic.
===================================================
2. TARGET AUDIENCE
===================================================
- Internal TV channel staff (editors, producers) - for quickly drafting a news piece.
- Potential partners/investors - as a technology demo.
===================================================
3. FUNCTIONAL REQUIREMENTS
===================================================
3.1 Homepage (presentation section)
- TV channel name, logo, brief project description.
- "How it works" block (3 steps: thesis - text - voiceover).
- "Try it" CTA button - leads to the interactive generation section.
3.2 News generation module (Gemini API)
- Language switcher for generation: Russian / Kazakh / English (affects both the Gemini prompt and the language of the resulting text).
- Thesis input field (text field, limit e.g. 200 characters).
- "Generate news text" button.
- Generation parameters (optional, user-selectable):
- Style: neutral / formal / conversational.
- Duration: 20 sec / 30 sec (affects text length, ~50-90 words per 20-30 sec of speech, adjusted per language).
- The generated text appears in an editable field - the user can manually adjust it before voiceover.
- Loading indicator while the API request is in progress.
- API error handling (timeout, rate limits, invalid response) - clear message shown to the user.
3.3 Voiceover module (ElevenLabs API)
- Voice selection from the list of ElevenLabs voices matching the selected language (male/female, voice name, preview if possible).
- Voiceover style/intonation selection (if supported by the ElevenLabs model - stability, expressiveness, etc., via voice_settings parameters).
- "Voice it" button.
- Loading/progress indicator for audio generation.
- In-page audio player (play/pause, seek, duration indicator).
- "Download audio" button - mp3 format only.
- API error handling.
3.4 Input moderation module (mandatory)
- Automatic check of the entered thesis before it is sent to Gemini, screening for:
- sexual content / explicit material;
- hate speech, discrimination;
- violence, extremism, insults;
- any other content unsuitable for broadcast.
- Implementation: Gemini's built-in safety settings PLUS an additional moderation layer on the backend (e.g., a separate moderation request or classifier) - do not rely on a single check alone.
- If disallowed content is detected, the request is not sent to Gemini, and the user sees the message "Blocked by moderation" (in the interface language - ru/kk/en).
- A similar check (or lack thereof) should also be considered for the final generated text before it is sent to TTS - a second check on output is recommended, since the model may rephrase text unpredictably.
- Log rejected requests for later review (no personal user data is stored, since there is no authentication).
3.5 History of generated news items (mandatory)
- Every generation (thesis, final text, language, selected voice/style, link to the mp3 file, creation date) is stored in a database on the backend.
- A separate "History" page/section listing past generations.
- Pagination of the list (e.g., 10-20 records per page, with page navigation or "load more").
- Ability to listen to/download audio directly from the history (without regenerating).
- The history is shared - a single feed of all generations, visible to all site visitors, not tied to a specific user/session (no authentication).
3.6 Optional (priority to be discussed)
- Ability to generate several text variants and pick the best one.
- Responsive layout (mobile version).
===================================================
4. NON-FUNCTIONAL REQUIREMENTS
===================================================
- Security: API keys (Gemini, ElevenLabs) are stored only on the backend and never exposed to the browser. All requests to external APIs go through a dedicated backend proxy.
- Performance: typical Gemini response time is 2-5 sec, ElevenLabs 3-10 sec depending on text length; the UI must clearly show a loading state.
- Rate limits: account for both APIs' quotas/rate limits; implement a queue or protection against spam requests (e.g., rate limiting per IP/session).
- Cross-browser support: latest versions of Chrome, Safari, Firefox, Edge.
- Responsiveness: desktop + mobile.
- Localization: interface and generation - Russian, Kazakh, English (language switcher on the page).
- Content moderation: mandatory on input (thesis) and recommended on output (generated text) - see section 3.4.
===================================================
5. USER FLOW
===================================================
Homepage
- Enter news thesis
- Generate text via Gemini (+ optional text editing)
- Select voice and style
- Generate voiceover via ElevenLabs
- Listen to / download audio
Openen op Upwork
AI proposal draft
Generate a short cover letter for this job. Edit before sending.
Sign in to generate an AI proposal draft.
Inloggen