Northquill.ai saas repair
Budget: $15.0 - $35.0
HOURLY / PART_TIME
⭐ 0.00 (0)
United States
artificial-intelligence, machine-learning, chatbot-development, natural-language-processing, python
Gewenste kwalificaties
- Ervaring: Gevorderd
Hi — I have an existing SaaS product called Northquill.ai.
Northquill takes an existing PowerPoint presentation plus narration and creates a finished narrated MP4.
The particularly important use case is NOT AI-generated narration. It is taking one existing continuous long-form recording and automatically determining when the presentation should advance through the original slides.
My real-world stress test is approximately:
- 4 hours of continuous existing narration
- roughly 150 PowerPoint slides
- no existing timestamp map
Northquill already has an AI alignment/review workflow, but it consistently fails at this scale.
The biggest problems are:
- suggested timings lean heavily toward approximately equal splitting of the total audio across the slides;
- the AI does not understand the semantic relationship between narration and individual slides well enough;
- some slides may legitimately remain on screen for seconds while others may remain for several minutes;
- the slide thumbnails in the existing review interface are currently failing to load;
- the review step exists, but too many of the AI recommendations need correction.
I do NOT want a simple improvement to the equal-duration heuristic or an LLM guessing timestamps.
I believe this is closer to a monotonic multimodal sequence-alignment problem.
A potential architecture would involve:
PowerPoint slide text / notes / visual concepts
+
WhisperX or equivalent word-level transcription / forced alignment
+
semantic matching between transcript sections and slides
+
high-confidence anchor detection
+
ordered/monotonic sequence alignment across the full deck
+
natural sentence / pause / silence boundary detection
+
confidence scoring
+
human review only for uncertain transitions
+
a deterministic timing map used for the final video render.
The original continuous audio should ideally remain intact. The output of the alignment system should primarily be accurate slide-change events rather than 150 separately cut audio files.
I am looking for someone to review the EXISTING Northquill codebase rather than rebuild the product from scratch.
For a first paid milestone, I would like to:
1. Review the current architecture and identify why its recommendations gravitate toward equal duration.
2. Diagnose the broken slide thumbnails/review UI.
3. Run the real ~4-hour / ~150-slide dataset through an improved alignment approach.
4. Compare the proposed slide boundaries to manually established correct transitions.
5. Produce a recommendation and proof of concept before making larger architectural changes.
The goal is that Northquill should do the hard alignment itself, with the human review step correcting exceptions rather than manually rebuilding the timeline.
Thewebsite northquill.ai is live and you can sign up for free and see the issues yourself. Im willing to share the code base separately
If you were approaching this problem, I would especially like to hear how you would solve the long-sequence semantic alignment portion rather than simply which AI APIs you would use.
Openen op Upwork
AI proposal draft
Generate a short cover letter for this job. Edit before sending.
Sign in to generate an AI proposal draft.
Inloggen