Developer Needed for Private AI Chat System (API-based)
Budżet: -
HOURLY / AS_NEEDED
⭐ 5.00 (2)
United Arab Emirates
python, api-integration, machine-learning, python-script, bot-development, data-extraction, api
Preferowane kwalifikacje
- Lokalizacja: United Arab Emirates
- Doświadczenie: Ekspert
- Angielski: Komunikatywny
- Job Success: 90%+
- Preferowany Rising Talent
Developer Needed for Private AI Chat System (API-based)
Project Overview
I need a private, self-hosted AI assistant that connects in real-time to third-party AI APIs (Anthropic Claude and/or Google Gemini) — NOT a local/offline LLM (no Ollama, no local model weights). The system must be accessible via a web interface from multiple devices (laptop, phone, tablet) and multiple locations, with my data staying under my control.
What I Need Built
1. Backend server that securely calls the Anthropic API and/or Google Gemini API
- API keys stored server-side only (never exposed to the browser/client)
- Support for switching between or combining providers
2. Simple web-based chat interface
- Responsive design (usable on desktop and mobile browsers)
- User login/authentication (so only I — and anyone I authorize — can access it)
3. Document storage & retrieval (RAG - Retrieval-Augmented Generation)
- Ability to upload my own documents/files (PDF, Word, text, etc.) through the web interface
- Documents automatically split into chunks and converted into embeddings (numerical representations for semantic search)
- Embeddings stored in a self-hosted vector database on my own server (e.g., Qdrant, Weaviate, or pgvector — developer to recommend based on scale)
- At query time, the system retrieves only the most relevant chunks (semantic similarity search) and sends those — not the full documents — to the AI API along with my question
- Full documents and the vector database must remain on my own infrastructure at all times; only the retrieved text snippets are sent externally to Claude/Gemini per query
- Ability to update/delete documents from the knowledge base (re-indexing when content changes)
- Source attribution: responses should indicate which document(s) the answer was drawn from
4. Content ingestion pipeline (for large personal knowledge sources)
- Ability to bulk-ingest a large personal library of trading education content (e.g., YouTube video transcripts) into the RAG knowledge base
- Pipeline should: fetch/extract transcripts, clean up the text (remove filler, fix formatting), chunk by topic/concept rather than fixed length, generate embeddings, and index into the vector database
- Should also support ingesting reference/technical documentation (e.g., the official Pine Script language reference) as a separate, always-available knowledge source used specifically to ground any code generation and reduce errors/hallucinations
- Reusable pipeline: I should be able to add new sources (new videos, new documents) later without needing the developer each time
5. Live web search capability
- The assistant should be able to search the web in real time for current information (e.g., market news, recent Pine Script/TradingView updates) when a question needs up-to-date data
- Implement via a web search tool/plugin connected to the AI API (e.g., Claude's built-in web search tool, or a search API such as Google/Bing/Perplexity Sonar integrated into the backend)
- Should be clearly distinguished from the RAG knowledge base: RAG = my own static documents/strategies; web search = live, current information from the internet
- Results should include source links so I can verify information
6. Hosting & deployment
- Deployed on a VPS (I will provide access — see below) or recommend one with justification
- Dockerized setup preferred, for portability and easy maintenance
- Must remain accessible 24/7 via a secure URL (HTTPS)
7. Logging & basic security
- Request logs (who asked what, when)
- Basic protection against unauthorized access (rate limiting, authentication)
8. Documentation
- Clear instructions on how to maintain, update, and restart the system
- How to add/rotate API keys
- How to add new users if needed
Requirements for the Developer
- Proven experience with backend development (Python/FastAPI or Node.js/Express)
- Experience integrating LLM APIs (Anthropic, OpenAI, or Google Gemini API)
- Experience with vector databases / RAG pipelines (e.g., Pinecone, Qdrant, Weaviate, or pgvector)
- Experience deploying and managing applications on a VPS using Docker
- Understanding of basic security practices (secrets management, authentication, HTTPS/SSL setup)
- Able to communicate clearly in English and explain technical choices in plain language
- Portfolio or examples of similar past projects (chatbots, AI integrations, RAG systems)
- Experience with text extraction/ingestion pipelines (e.g., transcript extraction, document parsing, bulk chunking strategies) is a plus
- Experience integrating web search tools/APIs (e.g., Claude's web search tool, Google Search API, Bing API, or Perplexity Sonar) is a plus
This will be a project based contract (Upwork fixed contract with milestones), where gradual payments will be released upon milestones achievements. It will be considered finalized when everything will be done/completed. This is not a "pay-by-the-hour" contract. The total budget & exact milestones for this project will be discussed and agreed upon.
Otwórz na Upwork
AI proposal draft
Generate a short cover letter for this job. Edit before sending.
Sign in to generate an AI proposal draft.
Zaloguj