← Jobb

Build and Integrate a Secure RAG System for Internal Company Documentation

Budget: $300.0 FIXED / ⭐ 0.00 (0) United States

machine-learning, artificial-intelligence, natural-language-processing, health-technology, chatbot-development, deep-learning, data-extraction, amazon-web-services, typescript, next.js, google-cloud-platform, react-js, pytorch, computer-vision

We are looking for an experienced AI/LLM engineer to design, build, and integrate a production-ready Retrieval-Augmented Generation (RAG) system for our company’s internal documentation. The system will allow employees to ask natural-language questions and receive accurate, context-aware answers based on our private technical and operational documents. Every answer should include references or links to the original source documents so users can verify the information. This is an urgent internal operations project. We are looking for someone who can quickly deliver a functional MVP while establishing an architecture that can be expanded over time. Current Documentation Our internal knowledge sources may include: PDF documents Microsoft Word documents Excel spreadsheets HTML pages and internal web documentation Product manuals and technical specifications SOPs and operational process documents Troubleshooting guides Bug and issue records Frequently updated vendor or product documentation The document collection contains both structured and unstructured data and may include multiple versions of the same document. Main Project Requirements 1. Document Ingestion Pipeline Build an automated document ingestion and processing pipeline that can: Upload and process PDF, Word, Excel, text, Markdown, and HTML files Extract text, tables, headings, document structure, and metadata Handle large technical documents Detect duplicate or updated documents Preserve the original document name, version, page number, section, URL, and other source metadata Support incremental updates without rebuilding the entire knowledge base Identify and report documents that failed to process The system should use an appropriate chunking strategy for technical documentation instead of relying only on fixed-length text splitting. 2. Vector Database and Retrieval Configure and integrate a vector database such as: Qdrant Pinecone Weaviate Milvus PostgreSQL with pgvector The retrieval system should support: Semantic vector search Keyword or BM25 search Hybrid retrieval Metadata filtering Product, document type, version, and category filtering Query rewriting or query expansion Reranking of retrieved results Retrieval across both English and Chinese documentation The engineer should recommend the most suitable embedding model and reranker for multilingual technical documentation. 3. RAG Answer Generation The system should: Generate answers using only retrieved internal information Include citations for every important claim Show the document name, section, page number, URL, or source location Avoid presenting unsupported information as fact Clearly state when the knowledge base does not contain enough information Support follow-up questions within the same conversation Maintain useful conversation context without allowing old conversation history to reduce retrieval accuracy Support configurable system prompts and response formats 4. Internal User Interface Build or integrate a simple internal interface where employees can: Ask questions View generated answers Open the cited source documents Review the retrieved source passages Filter searches by product, department, document type, or version Provide positive or negative feedback on answers Start a new conversation or review previous conversations A basic but functional internal web interface is acceptable for the MVP. 5. Administration and Knowledge-Base Management The system should provide an administrative workflow to: Upload documents Remove or replace documents Reprocess failed documents View ingestion status View document metadata Manage document categories Monitor retrieval and answer quality Review common unanswered questions Track user feedback 6. Security and Deployment Because the system will use private company documentation, security is important. Requirements include: Private deployment in our cloud environment or on-premises infrastructure Authentication and user access control Secure handling of API keys and credentials No unauthorized storage or use of company documents No use of our private documents to train public models Configurable document-level or role-based permissions Logging and auditability where appropriate Clear separation between development, testing, and production environments Please describe your recommended deployment approach and any external AI services that would receive company data. 7. Evaluation and Quality Testing The engineer should implement a practical evaluation process covering: Retrieval relevance Answer correctness Citation accuracy Hallucination rate Coverage of common internal questions Performance on multilingual queries Response latency Behavior when the answer is not present in the documentation The final delivery should include a test dataset or evaluation workflow that our team can continue using after the project is completed. Expected Deliverables The selected freelancer will deliver: A working RAG application deployed in our environment A document ingestion and synchronization pipeline Vector database configuration and indexing Hybrid retrieval and reranking implementation LLM answer-generation workflow with citations Internal search and chat interface Basic administrative document-management workflow Authentication and access-control implementation Evaluation results and testing documentation Architecture diagram Deployment instructions Source code with clear comments Environment configuration template Technical documentation and maintenance guide Knowledge-transfer session with our internal team A list of known limitations and recommended next steps Preferred Technical Stack We are open to recommendations, but relevant experience may include: Python FastAPI, Flask, or Django LangChain, LangGraph, LlamaIndex, or a custom RAG pipeline Qdrant, Pinecone, Weaviate, Milvus, or pgvector OpenAI, Azure OpenAI, Anthropic, Gemini, or privately hosted open-source models vLLM or other self-hosted model-serving frameworks Hugging Face embedding and reranking models React or Next.js Docker Kubernetes AWS, Azure, GCP, or on-premises deployment We value system quality and maintainability more than the use of any specific framework. Required Experience Applicants should have demonstrated experience with: Building production RAG systems Processing complex technical documents Vector databases and embedding models Hybrid search and reranking Citation and source-attribution systems LLM hallucination reduction Multilingual retrieval Secure enterprise or internal application deployment Docker-based deployment API and frontend integration RAG evaluation and performance optimization Nice-to-Have Experience Experience with GPU server or infrastructure documentation Experience with NVIDIA technical documentation Experience processing Excel-based product or BOM data Experience with knowledge graphs Experience with document versioning Experience with OCR and table extraction Experience deploying open-source LLMs through vLLM Experience integrating Jira, Confluence, SharePoint, Google Drive, or similar internal systems Experience building agent-based workflows around a RAG system
Öppna på Upwork