← İşler

Competitor Intelligence Engine Developer

Bütçe: - HOURLY / PART_TIME ⭐ 0.00 (0) Netherlands

agile-software-development, marketing, software-writing, python, sql, javascript, api, crawlers, postgresql, sqlalchemy, redis, data-extraction, api-integration

Tercih edilen nitelikler

  • Deneyim: Orta
  • İngilizce: Konuşma düzeyi
  • Job Success: 90%+
  • Rising Talent tercih edilir
Python developer — scale a competitor-monitoring engine (scraping, FastAPI, Apify) About us & the product: We're a Dutch online-marketing agency building an internal intelligence platform for our e-commerce clients. One module monitors the webshops of each client's competitors: it periodically snapshots their pages, diffs each snapshot against the previous one, and surfaces only meaningful changes as findings in an inbox — price moves, assortment changes, delivery-terms changes, positioning shifts, and CRO/conversion-element changes. Each finding gets an AI-written interpretation and concrete action points. This is not a greenfield project. The module is live in pilot and the hard parts exist: a pure, unit-tested diff engine with noise suppression (timestamps, stock counters, session IDs and A/B variants are stripped before comparison), deterministic structured-data extraction (JSON-LD → microdata → dataLayer → OpenGraph, merged per product, with an LLM fallback only for interpretive fields), idempotent change fingerprinting so re-runs never duplicate findings, multi-tenant row-level security, and cost tracking on every scrape and LLM call. The problem you'll solve: Today the scanner follows only ~7 hand-seeded pages per competitor, picked by arbitrary sitemap order — while these shops have catalogs of hundreds to thousands of products. Every page is fetched with a full browser render, one run at a time, which creates both a cost and a throughput ceiling. Our working plan: split monitoring into a broad cheap layer (whole catalog via batched plain-HTTP crawls + deterministic extraction, no LLM, stored as price/product time series) and a narrow expensive layer (a few key funnel pages via browser render + LLM, feeding the findings inbox). About that plan — and your role in it: The plan is worked out in detail, code-verified, and you'll receive the full document with acceptance criteria per step. But it was written by a small team and we know the difference between a plan and the truth: milestone 0 exists specifically to test its core assumptions against real data, and every milestone boundary is a point where we revise based on what we learned. We want a developer who challenges the plan where it's weak — if you see a better route, make the case with evidence and we'll change course. What we don't want is the opposite: silently building around something you disagree with. Disagreement goes on the table, before the code. One thing genuinely is settled: the stack. This is a production codebase, so we're not migrating frameworks or rewriting the engine — proposals need to work within it. Milestones (each independently shippable, fixed price per milestone; scope of later milestones is revisited as findings come in) M0 — Measurement day (1 day, paid trial). Run four defined measurements against our data and one pilot competitor — including batch-fetching 100 product URLs from a sitemap and reporting what percentage yields name/price/EAN through our existing extractor — plus a small, precisely specified config fix with extended tests. Explicitly a go/no-go: if the numbers contradict the plan, we adjust the plan. This milestone doubles as the paid trial for both sides. M1 — The broad layer (3–5 days). Scale an existing production-proven batch-crawl pattern to full catalogs: sitemap harvesting, batched cheerio crawls, deterministic extraction, per-hostname rate limiting and block-page detection, crawl budgets per client tier, provenance-flagging so auto-discovered products don't pollute hand-curated data. M2 — Smarter page selection (1–2 days). Deterministic ranking instead of "first 3 in sitemap order", plus audit events when discovery silently finds nothing. M3 — Sitemap as change oracle (2–3 days). Parse lastmod + full URL sets; URL-set diffing as a whole-catalog assortment signal. Additive only, with sanity checks — sitemaps lie. M4 — CRO detection (3–4 days). On HTML we already fetch: element counting (form fields, payment icons, trust badges), a normalized DOM-skeleton hash for restructure detection, A/B-test detection via repeated sampling with pinned proxy geo, urgency-signal normalization. M5 — Throughput & monitoring (1–2 days). Higher concurrency for the cheap layer, scheduling that fits ~40 clients, alerting on overruns. Stack: Python 3.12, FastAPI, SQLAlchemy 2.0 async, Alembic, Postgres with RLS, Redis + RQ, APScheduler, Apify (cheerio & Playwright crawlers), Anthropic API behind provider interfaces, pytest, GitHub Actions. Conventions: boring over clever, tests ship with the change, small PRs per milestone. What we provide: repo access, the detailed plan document with code-level references, scraping budget for dry-runs, and a responsive technical owner (the codebase author) for questions and same-day review. One rule has a story behind it: an error page was once interpreted as "competitor removed everything", so production baselines are sacred here — changes near the diff engine need a written case first, never a quiet workaround. Who we're looking for: strong async Python with real production scraping experience — you know why JSON-LD beats LLM extraction on product pages, what a 200-OK block page is, and why cheerio-vs-headless is a cost decision, not a preference. You push back with arguments when you disagree, commit small and push daily, flag blockers early, and you're transparent about how you work — including which AI coding tools you use and where you draw the line. Dutch is not required; the target pages are Dutch-language webshops, so comfort with non-English content matters. Sizing: 11–17 working days total. M0 first as the paid trial; continuation per milestone after review.
Upwork'te aç

AI proposal draft

Generate a short cover letter for this job. Edit before sending.

Sign in to generate an AI proposal draft.

Giriş yap