AI/Python Developer Needed – Build a Local RAG Research Assistant
Бюджет: $350.0
FIXED /
⭐ 5.00 (13)
USA
python
AI/Python Developer Needed – Build a Local RAG Research Assistant
We are looking for an experienced Python/AI Developer to build a local Retrieval-Augmented
Generation (RAG) research assistant for a single user. This is a personal research tool, not an
enterprise application.
The objective is to create an AI assistant that can answer questions by searching a knowledge base of
300–500 research papers (PDFs) using semantic search and a Large Language Model (LLM). Every
response should be based on retrieved documents and include citations to the original source.
Scope of Work
The selected freelancer will be responsible for the complete development of the application.
1. Research Paper Collection
• Source and curate 300–500 publicly available research papers.
• Preferred domain: Cybersecurity (network security, malware analysis, threat intelligence, cloud
security, zero trust, vulnerability management, etc.).
• If you believe another technical domain (AI, Machine Learning, Medical Research, etc.) offers a
stronger public dataset, explain why in your proposal.
2. Document Ingestion Pipeline
Develop an automated pipeline that can:
• Import PDF files.
• Extract text from PDFs.
• Clean and preprocess text.
• Split documents into chunks.
• Generate embeddings.
• Store embeddings in a vector database.
• Allow additional PDFs to be added later without rebuilding the entire database.
3. Vector Database
Use one of the following:
• ChromaDB (preferred)
• FAISS
• Qdrant
Explain your choice if using another vector database.
4. RAG Pipeline
Build the complete Retrieval-Augmented Generation workflow using Python.
The solution may use:
• LangChain
• LlamaIndex
• A custom implementation
The pipeline should:
• Retrieve the most relevant document chunks.
• Send retrieved context to the LLM.
• Generate answers based only on retrieved information.
• Minimize hallucinations.
• Return citations for every answer.
5. LLM Integration
Integrate one of the following:
• OpenAI GPT
• Anthropic Claude
• Gemini
• Local open-source model
Please explain which model you recommend and why.
6. User Interface
Build a simple web interface using Streamlit or Gradio.
The interface should allow the user to:
• Ask questions in natural language.
• View AI-generated answers.
• See the source documents used.
• View page numbers (where available).
• Continue asking follow-up questions within the same conversation.
Required Features
The completed application must support:
• Semantic search
• Natural language question answering
• Multi-turn conversations
• Document summarization
• Citation-aware responses
• Source document references
• Fast retrieval
• Local execution
• Easy addition of new research papers
Deliverables
The completed project must include:
• Complete Python source code
• PDF ingestion pipeline
• Vector database
• Working RAG application
• Streamlit or Gradio interface
• Installation guide
• README documentation
• Requirements file
• Instructions for updating the knowledge base
• Basic testing to verify functionality
Required Skills
• Python
• RAG (Retrieval-Augmented Generation)
• LangChain or LlamaIndex
• OpenAI API (or equivalent)
• Vector Databases
• ChromaDB / FAISS / Qdrant
• NLP
• Semantic Search
• PDF Processing
When Applying
Please include:
1. Links to previous RAG or LLM projects.
2. The technology stack you recommend.
3. Your proposed architecture.
4. Estimated timeline.
5. Fixed-price quote.
Important: Generic AI proposals will be ignored. Please explain how you would implement this project,
including your preferred RAG framework, embedding model, vector database, and LLM.
Открыть заказ