← Live-лента

Senior Python Engineer for Distributed Parallel Pipeline (Celery + Redis + Flower)

Бюджет: $30.0 - $85.0 HOURLY / PART_TIME ⭐ 0.00 (0) United States

redis, python, celery, mpd

Предпочтительная квалификация

  • Локация: Sri Lanka
  • Опыт: Эксперт
About the work We run an AI-powered data ingestion platform. Our Python pipeline crawls, extracts, validates, and scores large volumes of records, and it is live in production today. The pipeline runs on Celery with Redis as the broker and Flower for monitoring, all three working together as one parallel processing system. Please read this before applying This is not a beginner role and it is not a learning opportunity. We need an engineer who has already built and debugged distributed parallel pipelines in production. If this would be your first time doing this kind of work, this is not the right project for you. We also need to be direct about AI. We are an AI company and we use AI tooling ourselves, so this is not about the tools. It is about expertise. We are not hiring someone who will lean on an AI assistant to teach them Celery concurrency and task routing while working on our production system. We have already been through one engagement where the parallel work did not hold up, and we are not repeating it. The knowledge has to be yours before you start. Someone who has used Celery to send background emails, or Redis as a cache, is not who we are looking for. We need someone who has designed task fan-out at volume and knows why a pipeline that looks parallel on paper runs like it is sequential. The problem we need solved first Our parallel implementation is not producing a speedup. End-to-end runtimes are roughly the same as the single-threaded version it replaced. Workers are running, tasks are being dispatched, Flower shows activity, and yet the throughput gain is not there. This does not make sense to our team and we need an expert to diagnose it and fix it. The engineer who originally built the parallel layer is no longer with us, so you would be coming into existing code and forming your own read on it. Phase two, after the above is resolved There is a structural inefficiency in the code where each URL is fetched and processed twice across two stages, when a single fetch could serve both. Collapsing that into one pass is a separate piece of work we want to scope after the parallelism issue is solved. Mentioning it here so you understand where the project is going. How this engagement works 1. Paid 30-minute consultation. We walk you through the current architecture and observed runtimes. You give us your initial read and your rough sense of effort and cost. 2. If it is a fit, we scope and award the diagnostic and fix work as a separate contract. 3. Phase two, and potentially ongoing pipeline work, follows from there. We are starting with the paid consultation because we want to make sure we are aligned before anyone touches the codebase. Required experience ∙ Production experience with Celery, Redis, and Flower operating together in a parallel or distributed pipeline, not each used separately in unrelated projects ∙ Hands-on work with task fan-out patterns: groups, chords, chunking, task routing, and dedicated queues ∙ Experience tuning worker concurrency, pool types (prefork, gevent, eventlet, threads), prefetch multiplier, and acknowledgment settings ∙ Track record diagnosing pipelines where the parallelism failed to produce a throughput gain ∙ High-volume web crawling or scraping at scale, including rate limiting, retries, and failure handling ∙ Strong Python, and the judgment to tell us when our current approach is wrong Nice to have ∙ Experience with pipelines that call LLM or other paid APIs per record, where cost per record matters as much as speed ∙ MariaDB or MySQL experience at volume ∙ Experience running Celery on self-managed Linux infrastructure What to include in your proposal Please do not send a generic or AI-generated proposal. We are looking for evidence of work you have actually done, so include as much of the following as you can: ∙ Two or three parallel pipelines you have personally built or fixed. For each, tell us what the work was, how you distributed it, and what the before and after numbers looked like. Records or tasks per hour, total runtime, whatever metric you used. ∙ Scale details: how many workers, what concurrency settings, what infrastructure, what the broker was handling. ∙ Whether you built the pipeline from scratch or inherited and tuned someone else’s code. Both are relevant to us, but we want to know which. ∙ Links to public repos, architecture write-ups, or redacted diagrams if you have them. Screenshots of Flower or monitoring dashboards from past work are welcome. ∙ Upwork clients or references from similar distributed pipeline work. ∙ Your time zone and your overlap with US Eastern. Proposals without specific pipeline examples will not get a response. Expect to discuss your past work in detail on the call. Logistics ∙ NDA and IP assignment required before any code or environment access
Открыть заказ

AI-черновик отклика

Короткий текст отклика для копирования в оффер: интерес + готовность работать.

Войдите, чтобы сгенерировать AI-черновик.

Войти