← Zakázky

Senior Python Engineer for Financial Data Engineering Microservice

Rozpočet: $80.0 - $120.0 HOURLY / FULL_TIME ⭐ 4.92 (85) United States

microservices, postgresql, redis, software-architecture, api-development, restful-api, business-logic-layer, database-architecture, sql, web-services, database-design, python, performance-tuning

Preferred qualifications

  • Experience: Expert
  • English: Fluent
  • Job Success: 90%+
  • Rising Talent preferred
  • Min. earnings: $10,000+
Summary We are commissioning a high-performance bank attribute calculation microservice in Python that ingests normalized open-banking data (Plaid item, accounts, and transaction history) for a single applicant and computes a large, deterministic set of cash-flow and risk attributes that drive our automated, subprime small business underwriting engine. The framework scales to hundreds of thousands of computed attributes per applicant, of which a curated subset (~20,000) is returned synchronously to the caller. We are looking for a Senior Python Engineer / contractor to architect and deliver this service end-to-end, with correctness, a hard latency SLA, and clean, extensible system design as the top priorities. Downstream consumption of the resulting feature contract is intentionally agnostic — the service should not be designed toward, or coupled to, any specific downstream model or consumer. Core Responsibilities Ingestion & Normalization: Parse and normalize noisy, nested Plaid-style JSON (item, accounts, transactions) into a clean internal schema, correctly handling the Plaid amount sign convention, nested sub-objects, pending transactions, and missing/null fields. Deduplication & Account Resolution: Deduplicate transactions arising from the same physical account being connected multiple times, using a tiered account-identity key (persistent account ID first, falling back to an institution/subtype/mask composite and its name-based variants). Transfer Detection & Balance Dating: Detect inter-account transfers between an applicant's own accounts (amount and date-window matching) and derive a canonical balance date with an authorized-date-first, transaction-date-fallback rule. Merchant Enrichment & Classification: Integrate two static merchant-resolution lookup tables (merchant name and merchant less-transaction fallback) and port deterministic, tiered regex/keyword rules for income and expense classification, preserving tier precedence and exclusion logic exactly. Balance Reconstruction: Reconstruct a 91-day backward daily balance time series per applicant from an anchor balance and net daily flows, producing account-level and applicant-level daily rollups. Attribute Computation Engine: Design and scale a highly extensible, declaratively configurable rules and math framework computing statistics across sliding/cumulative time windows and multiple dimensions (category, merchant, PFC, and cross-dimension combinations), additionally sliceable by cross-cutting axes such as payment channel, at 100k+ feature scale — where adding a new window, dimension, statistic, or slicing axis is a configuration change, not a rewrite. Curated Subset Design: Own the actual selection methodology for the ~20,000-feature curated response — a deterministic, documented approach (e.g., by stability, coverage, or business relevance), not a naive random sample of feature names. Persistence Architecture: Design a persistence layer that synchronously and atomically records the raw input, the curated feature set, and run/version metadata as the audit record, while writing the full feature set off the request path as a re-computable/derived projection (e.g., compressed cache write plus background object storage upload). Performance & SLA Engineering: Meet a hard end-to-end latency SLA (see below) via startup-time (not per request) compilation of the feature computation graph, vectorized/parallelized execution, and a memory stable design under concurrent load — including when the feature surface is multiplied by slicing axes like payment channel. API Development: Design and maintain a clean, versioned, self-describing request/response contract for payload submission and curated feature retrieval. Key Constraints & Acceptance Criteria Latency: End-to-end (ingestion through full computation, synchronous persistence, and response) must be under 3 seconds, with documented p99 benchmarks at target concurrency. Determinism: Identical inputs must always produce identical outputs, computation must be hermetic (no live external lookups — static lookups are versioned inputs loaded at startup), and every run must record the exact service version for faithful historical re-computation. Reference Parity: Computed attributes must match known reference examples exactly, subject to a documented floating-point rounding policy. Scale & Stability: The service must remain memory-stable under high-concurrency load despite a 100k+ feature space, with vectorized/parallelized computation demonstrated via benchmarks. Our Technical Stack & Environment The architecture and technology stack are largely the contractor's decision. The binding constraints are that the service must run on AWS and be deployable via a Docker container (we use EKS). Our platform team owns core DevOps, deployment pipelines, and infrastructure scaling, so you can focus on application architecture, performance, and features. For context only (non-binding — you may adopt, substitute, or replace any of this): our current environment uses Python 3.11+, FastAPI, Pydantic, AsyncIO, columnar/vectorized libraries (Polars / NumPy / Pandas), PostgreSQL, and Redis. If your recommended stack diverges (e.g., a different cache/queue technology), we ask for a resourcing/TCO estimate as soon as you're reasonably confident in that choice. Ideal Candidate Profile • Senior Backend/Data Architect: Proven track record designing data-intensive backend microservices from scratch, including persistence and caching architecture for high-volume-derived data. • Fintech or Risk Fluent: Deep experience with transactional financial datasets, financial feature engineering, or risk/underwriting systems; comfortable porting SQL/notebook research logic (rule ladders, exclusions, tiered classification) into production code faithfully. • Performance-Driven: Skilled at wringing performance out of Python — vectorized/columnar computation, async patterns, and avoiding memory bottlenecks at large feature-space scale. • Obsessed with Correctness: Understands that in credit risk, data correctness is non-negotiable; writes highly testable, deterministic code with rigorous edge-case handling (null dates, missing identifiers, sparse histories, negative/corrupted balances). Key Deliverables • Production-ready Python microservice implementing the full ingestion → cleaning/enrichment → balance reconstruction → feature computation → persistence → curated-response pipeline, deployable via Docker on AWS. • A scalable, extensible attribute-calculation framework where new windows, dimensions, and statistics are additive/configurable. • A persistence design meeting the atomic-audit-record and re-computable-full-feature-set guarantees, including crash-recovery behavior. • Comprehensive unit and integration test suites proving computation logic matches known reference examples, plus edge-case coverage. • Performance benchmarks verifying the latency SLA and memory stability under concurrency. • Clear documentation of the request, internal, lookup, and curated-response contracts.
Otevřít na Upwork

AI proposal draft

Generate a short cover letter for this job. Edit before sending.

Sign in to generate an AI proposal draft.

Přihlásit