Python Real-Time Audio + AI Integration Developer – Windows Desktop App
Budget: $75.0 - $150.0
HOURLY / PART_TIME
⭐ 0.00 (0)
United States
python, websockets, api-integration, artificial-intelligence, audio-engineering
Bevorzugte Qualifikationen
- Erfahrung: Experte
Project: Real-Time Audio Protection Tool for Live Streamers
I need a Python desktop application built for Windows 11 that runs silently in the background and protects a live streamer from accidental platform violations in real time across TikTok, YouTube, and Facebook simultaneously.
What the app does:
The app captures live microphone audio, streams it to Deepgram Nova-3 via WebSocket for real-time speech transcription, checks every word against a per-platform banned word list, runs a local Ollama Llama 3.3 8B model for contextual violation detection, and when either detection layer fires it replaces that specific audio window with clean silence before it ever reaches OBS or any streaming platform. No bleep sound. Just a natural silent pause. The streamer's Kick, Rumble, and Locals streams run completely untouched with no filtering applied.
Exact technical requirements:
-Python 3.11 on Windows 11
-PyAudio for real-time microphone capture
-Deepgram Nova-3 WebSocket SDK for sub-300ms streaming transcription
-Local Ollama Llama 3.3 8B for contextual moderation via rolling 30 second transcript window with 150ms hard timeout
-VB-Cable virtual audio routing so OBS receives the clean filtered audio
-Per-platform word lists for TikTok, YouTube, and Facebook as core platforms with Kick and Rumble as optional toggles the user can switch on
-Silent mute gate using a 1 second audio buffer, no bleep sounds
-Simple one-button Tkinter GUI, Start Protection and Stop Protection, designed for a non-technical user
-Single .bat launcher file so the end user just double clicks to run it
-Clean commented code so word lists can be updated easily when platform policies change
What I need from you:
A fully working application I can install on one Windows 11 machine, test end to end, and hand to a non-technical streamer who will only ever see the one button. The code needs to be clean, documented, and structured so I can update the banned word lists myself without touching the rest of the application.
What you must have to be considered:
-Demonstrated experience with PyAudio or sounddevice in a real shipped project, share a GitHub link or code sample showing real-time audio work
-Experience integrating a streaming STT API, Deepgram SDK is a strong bonus, any real-time streaming STT counts, Whisper only experience does not qualify
-Experience running local LLMs via Ollama or llama.cpp
-Knowledge of Windows virtual audio routing and VB-Cable
-Experience building Python desktop GUIs with Tkinter or PyQt
Do not apply if:
-You have only worked with batch transcription and not live streaming audio
-You want to build a web app or browser extension instead of a Windows desktop application
-You have no real-time audio pipeline experience
Budget: $75 to $150 per hour
Timeline: 40 to 60 hours for a complete working build
To apply: Send me a brief explanation of how you would structure the audio pipeline specifically, and include any GitHub links or code samples showing real-time audio or STT API work. Applications without this will not be reviewed.
Auf Upwork öffnen
AI proposal draft
Generate a short cover letter for this job. Edit before sending.
Sign in to generate an AI proposal draft.
Anmelden