ML Engineer / Data Scientist: Retail Demand Forecasting Precision Improvement
Budget: $800.0
FIXED /
⭐ 0.00 (0)
United Kingdom
machine-learning, python-sklearn, statistics
Type: Contract, phased (diagnostic first, build-out optional)
Company: Arka Global (ShopSmartr platform)
About the project
ShopSmartr is a restock automation platform used in live UK retail stores. It reads historical purchase invoices and POS barcode sales, learns demand patterns per product, and automatically adds recommended restock quantities to a supplier cart, no manual invoice or wastage entry by the store manager.
The forecasting model is currently in production and running at roughly 85% accuracy. That gap is large enough that store managers still have to manually review and correct the cart, which undermines the "hands-off" value of the bot. We're looking for someone to diagnose where the accuracy ceiling is coming from and improve it.
What we need
Phase 1 (paid diagnostic, ~1 week, fixed scope):
Review our current pipeline, feature set, and model approach
Review a sample of our historical invoice and POS sales data
Identify the actual sources of error (data quality vs. model architecture vs. feature gaps vs. inherent demand noise)
Give us a realistic accuracy/precision target and a plan to get there, with an estimate of effort
Phase 2 (build-out, scope and cost set after Phase 1):
Implement the improvements agreed in Phase 1
Help us decide whether to optimize primarily for precision (avoiding wrong items landing in the cart) or recall (avoiding missed restocks), and tune accordingly
Document the approach so our internal team can maintain it
Known context (please read before applying)
Product barcodes are sometimes discontinued and reissued when a product's price changes, which breaks time-series continuity for that SKU in the training data. We suspect this is part of the accuracy gap. Experience with this kind of data quality issue in retail/POS datasets is a strong plus.
The pipeline already includes skip rules, tiered quantity logic (top/medium/slow sellers), shelf reserves, and supplier-specific ordering rules. The forecasting model sits underneath this, so changes need to integrate with existing logic, not replace it.
Data is proprietary; a sanitized/sample dataset will be provided for Phase 1, full production data access requires an NDA.
What we're looking for
Proven experience with demand forecasting specifically (retail, FMCG, or similar), not just general ML
Comfortable with intermittent/sparse demand series, seasonality, and promo effects
Python (pandas, scikit-learn, or similar); experience with time-series specific libraries a plus
Able to work with messy real-world POS/invoice data and call out data quality issues, not just model issues
To apply, please include
A brief example of a similar forecasting problem you've solved (industry, what accuracy/precision you achieved, and from what baseline)
Your approach to diagnosing whether an accuracy ceiling is a data problem or a model problem
Open job