Python Backend / AI Platform Engineer — Data Pipelines, RAG, Agents & AWS
Бюджет: $17.0 - $25.0
HOURLY / FULL_TIME
⭐ 5.00 (2)
United States
data-cleansing, data-segmentation
Бажана кваліфікація
- Досвід: Експерт
We are looking for an individual Python backend engineer to join an existing team building an internal AI knowledge platform.
Your role is to take primary ownership of the backend systems that make an AI application reliable: document ingestion, data pipelines, retrieval/RAG, AI agents and tools, integrations, background processing, and AWS infrastructure.
What you will work on
Document and data pipelines
- Ingest PDF, DOCX, XLSX, CSV, and other business documents
- Reliably extract and normalize content
- Build chunking, metadata, embedding, and indexing pipelines
- Handle large files, retries, failures, reprocessing, and partial processing
- Make ingestion idempotent and prevent duplicate data
- Debug corrupted or incomplete document extraction
- Build synchronization with external sources such as Google Drive
RAG and retrieval
- PostgreSQL + pgvector
- Embeddings and semantic retrieval
- Metadata filtering
- Permission-aware retrieval
- Context assembly
- Retrieval quality and latency
- Debugging incorrect or incomplete AI answers
- Evaluation and observability
AI agents and tools
We are beginning to add agents and tools to the application.
Initial capabilities include:
- Web search
- Internal knowledge search
- External API calls
- File generation
You will help build the shared architecture behind these capabilities, including:
- Tool/function calling
- Agent orchestration
- Tool schemas
- Context and state management
- Permissions
- Timeouts and retries
- Error handling
- Logging/tracing
- Reusable patterns that make additional tools easy to add
We do not want every agent implemented as a separate one-off system.
AWS and infrastructure
Our application runs primarily on AWS.
Relevant technologies include:
- ECS / Fargate
- RDS PostgreSQL
- S3
- Redis / ElastiCache
- IAM
- Secrets Manager
- CloudWatch
- AWS Bedrock
- Docker
- GitHub Actions / CI-CD
Current stack
- Python
- FastAPI
- PostgreSQL / pgvector
- Celery
- Redis
- AWS ECS Fargate
- AWS RDS
- AWS S3
- AWS Bedrock
- Docker
- GitHub Actions
- Angular / PrimeNG frontend
You do NOT need to be an Angular specialist. Another developer handles most frontend work.
The person we want
Your strongest area should be Python backend/data engineering.
We are particularly interested in engineers who have built systems where data moves through multiple processing stages and failures must be handled correctly.
You should be comfortable investigating problems such as:
- Why was the same file processed twice?
- Why did extraction work for most PDFs but fail for one?
- How should an ingestion job be made idempotent?
- How should modified Google Drive files be detected and reprocessed?
- What happens if a worker dies halfway through processing?
- Why did a RAG system retrieve the wrong information?
- Was a bad answer caused by retrieval, context assembly, or the LLM?
- How should an agent tool fail without breaking the conversation?
Required experience
Strong experience with:
- Python backend development
- FastAPI or another Python API framework
- PostgreSQL
- Background/asynchronous processing
- Data ingestion, ETL, or synchronization pipelines
- APIs
- Docker
- AWS
- Debugging production systems
Experience with several of the following is strongly preferred:
- pgvector or another vector database
- Production RAG systems
- Celery / Redis
- Document parsing
- Google Drive API
- Incremental synchronization
- LLM function/tool calling
- AI agent architectures
- AWS Bedrock
- ECS / Fargate
- CI/CD
This is probably not a fit if
- Your background is primarily frontend development
- Most of your AI experience is building basic chatbot demos
- Your primary experience is connecting LangChain to OpenAI APIs
- You have not worked on production data pipelines
- You are an agency submitting a developer
- You will not personally be doing the work
How we work
You will work with an existing developer and technical/product lead.
We expect you to:
- Understand the existing system before rewriting things
- Find root causes rather than repeatedly patching symptoms
- Make small, reviewable changes
- Explain architectural tradeoffs
- Write maintainable production code
- Add appropriate logging and observability
- Document important infrastructure and architectural decisions
At least two hours of overlap with US Eastern business hours is required.
Applying
Do NOT send us a long cover letter.
Your cover letter must be under 100 words. Do not summarize the job description or explain why you are passionate about data pipelines, file processing, and agents.
We will primarily evaluate you based on:
1. Your answers to the screening questions
2. Whether you followed all application instructions
3. Relevant past work
4. A short technical interview
5. A small paid technical trial for finalists
Please review the attachment before submitting your proposal.
We are looking for someone who can become an ongoing member of the development team if the initial engagement goes well.
Відкрити замовлення
AI-чернетка відгуку
Короткий текст відгуку для копіювання в офер: інтерес + готовність працювати.
Увійдіть, щоб згенерувати AI-чернетку.
Увійти