Senior Python Engineer for Financial Data Engineering Microservice
Orçamento: $80.0 - $120.0
HOURLY / FULL_TIME
⭐ 4.92 (85)
United States
microservices, postgresql, redis, software-architecture, api-development, restful-api, business-logic-layer, database-architecture, sql, web-services, database-design, python, performance-tuning
Qualificações preferidas
- Experiência: Especialista
- Inglês: Fluente
- Job Success: 90%+
- Rising Talent preferido
- Ganhos mín.: $10,000+
Summary
We are commissioning a high-performance bank attribute calculation microservice in Python that ingests normalized open-banking data (Plaid item, accounts, and transaction history) for a single applicant and computes a large, deterministic set of cash-flow and risk attributes that drive our automated, subprime small business underwriting engine.
The framework scales to hundreds of thousands of computed attributes per applicant, of which a curated subset (~20,000) is returned synchronously to the caller. We are looking for a Senior Python Engineer / contractor to architect and deliver this service end-to-end, with correctness, a hard latency SLA, and clean, extensible system design as the top priorities. Downstream consumption of the resulting feature contract is intentionally agnostic — the service should not be designed toward, or coupled to, any specific downstream model or consumer.
Core Responsibilities
Ingestion & Normalization: Parse and normalize noisy, nested Plaid-style JSON (item, accounts, transactions) into a clean internal schema, correctly handling the Plaid amount sign convention, nested sub-objects, pending transactions, and missing/null fields.
Deduplication & Account Resolution: Deduplicate transactions arising from the same physical account being connected multiple times, using a tiered account-identity key (persistent account ID first, falling back to an institution/subtype/mask composite and its name-based variants).
Transfer Detection & Balance Dating: Detect inter-account transfers between an applicant's own accounts (amount and date-window matching) and derive a canonical balance date with an authorized-date-first, transaction-date-fallback rule.
Merchant Enrichment & Classification: Integrate two static merchant-resolution lookup tables (merchant name and merchant less-transaction fallback) and port deterministic, tiered regex/keyword rules for income and expense classification, preserving tier precedence and exclusion logic exactly.
Balance Reconstruction: Reconstruct a 91-day backward daily balance time series per applicant from an anchor balance and net daily flows, producing account-level and applicant-level daily rollups.
Attribute Computation Engine: Design and scale a highly extensible, declaratively configurable rules and math framework computing statistics across sliding/cumulative time windows and multiple dimensions (category, merchant, PFC, and cross-dimension combinations), additionally sliceable by cross-cutting axes such as payment channel, at 100k+ feature scale — where adding a new window, dimension, statistic, or slicing axis is a configuration change, not a rewrite.
Curated Subset Design: Own the actual selection methodology for the ~20,000-feature curated response — a deterministic, documented approach (e.g., by stability, coverage, or business relevance), not a naive random sample of feature names.
Persistence Architecture: Design a persistence layer that synchronously and atomically records the raw input, the curated feature set, and run/version metadata as the audit record, while writing the full feature set off the request path as a re-computable/derived projection (e.g., compressed cache write plus background object storage upload).
Performance & SLA Engineering: Meet a hard end-to-end latency SLA (see below) via startup-time (not per request) compilation of the feature computation graph, vectorized/parallelized execution, and a memory stable design under concurrent load — including when the feature surface is multiplied by slicing axes like payment channel.
API Development: Design and maintain a clean, versioned, self-describing request/response contract for payload submission and curated feature retrieval.
Key Constraints & Acceptance Criteria
Latency: End-to-end (ingestion through full computation, synchronous persistence, and response) must be under 3 seconds, with documented p99 benchmarks at target concurrency.
Determinism: Identical inputs must always produce identical outputs, computation must be hermetic (no live external lookups — static lookups are versioned inputs loaded at startup), and every run must record the exact service version for faithful historical re-computation.
Reference Parity: Computed attributes must match known reference examples exactly, subject to a documented floating-point rounding policy.
Scale & Stability: The service must remain memory-stable under high-concurrency load despite a 100k+ feature space, with vectorized/parallelized computation demonstrated via benchmarks.
Our Technical Stack & Environment
The architecture and technology stack are largely the contractor's decision. The binding constraints are that the service must run on AWS and be deployable via a Docker container (we use EKS). Our platform team owns core DevOps, deployment pipelines, and infrastructure scaling, so you can focus on application architecture, performance, and features.
For context only (non-binding — you may adopt, substitute, or replace any of this): our current environment uses Python 3.11+, FastAPI, Pydantic, AsyncIO, columnar/vectorized libraries (Polars / NumPy / Pandas), PostgreSQL, and Redis. If your recommended stack diverges (e.g., a different cache/queue technology), we ask for a resourcing/TCO estimate as soon as you're reasonably confident in that choice.
Ideal Candidate Profile
• Senior Backend/Data Architect: Proven track record designing data-intensive backend microservices from scratch, including persistence and caching architecture for high-volume-derived data.
• Fintech or Risk Fluent: Deep experience with transactional financial datasets, financial feature engineering, or risk/underwriting systems; comfortable porting SQL/notebook research logic (rule ladders, exclusions, tiered classification) into production code faithfully.
• Performance-Driven: Skilled at wringing performance out of Python — vectorized/columnar computation, async patterns, and avoiding memory bottlenecks at large feature-space scale.
• Obsessed with Correctness: Understands that in credit risk, data correctness is non-negotiable; writes highly testable, deterministic code with rigorous edge-case handling (null dates, missing identifiers, sparse histories, negative/corrupted balances).
Key Deliverables
• Production-ready Python microservice implementing the full ingestion → cleaning/enrichment → balance reconstruction → feature computation → persistence → curated-response pipeline, deployable via Docker on AWS.
• A scalable, extensible attribute-calculation framework where new windows, dimensions, and statistics are additive/configurable.
• A persistence design meeting the atomic-audit-record and re-computable-full-feature-set guarantees, including crash-recovery behavior.
• Comprehensive unit and integration test suites proving computation logic matches known reference examples, plus edge-case coverage.
• Performance benchmarks verifying the latency SLA and memory stability under concurrency.
• Clear documentation of the request, internal, lookup, and curated-response contracts.
Abrir na Upwork
AI proposal draft
Generate a short cover letter for this job. Edit before sending.
Sign in to generate an AI proposal draft.
Entrar