← Joburi

AWS/Python Data Engineer — Secure Student Data Storage + Pseudonymization MVP

Buget: $80.0 - $150.0 HOURLY / PART_TIME ⭐ 4.93 (12) United States

amazon-web-services, amazon-s3, python

Unlock Education is an education analytics company working with U.S. public schools and districts. We analyze student-level assessment and instructional data and need to establish a lightweight, secure workflow for handling FERPA-protected student records. We already have an active AWS environment. We are not looking to build a new cloud platform or data lake. We need an experienced AWS/Python engineer to extend our existing environment with a small, well-designed secure data area and a simple pseudonymization utility. Our immediate goal is: District secure SharePoint → restricted AWS storage → pseudonymization → analysis-ready files The engagement should take approximately 5–10 hours and leave us with a simple system we can operate ourselves. SCOPE 1. Review our existing AWS environment Briefly review the current AWS account/configuration and determine the cleanest way to add secure storage for district student data without disrupting our existing application infrastructure. We expect this will primarily involve S3, IAM and appropriate encryption/logging rather than new application infrastructure. 2. Create three access-controlled data areas Set up logically and technically separated storage for: Raw / Restricted — original district files as received, potentially containing district student IDs, Florida Student IDs and other direct identifiers. Identity / Crosswalk — mapping between district identifiers and randomly generated Unlock analytical IDs. This should have the most restrictive access. Analysis — pseudonymized datasets containing only the fields needed for analysis. Access should be technically enforced through AWS IAM. Routine analysts should be able to access Analysis without access to Raw or Identity. We want a simple implementation appropriate to our current scale—not enterprise infrastructure. 3. Build a lightweight Python pseudonymization/preprocessing utility Create a documented Python script that: • reads CSV/XLSX district files; • identifies district student IDs; • generates a cryptographically random opaque Unlock student ID for each new student; • maintains the same Unlock ID for that student across files and subsequent runs; • maintains the district-ID ↔ Unlock-ID crosswalk in the restricted Identity area; • replaces district student IDs with Unlock IDs in analytical outputs; • removes configured direct identifiers such as student name, district ID, Florida Student ID, DOB, email, etc.; • optionally supports the same approach for teacher and section identifiers; • writes pseudonymized files to the Analysis area. We do not want student tokens generated using a simple deterministic hash of the original ID. The script should be configuration-driven enough that we can specify which columns are identifiers/removable fields without rewriting the code for every district. 4. Add basic validation/QA The process should report basic checks such as: • input/output row counts; • unique student counts; • new versus existing IDs mapped; • missing/null student IDs; • duplicate or problematic IDs; • removed fields; • output files generated. The purpose is simply to give us confidence that pseudonymization has not broken joins or unexpectedly altered the analytical data. 5. Test using synthetic data The contractor should not require access to actual student records. We will provide synthetic files approximating our district schemas. The solution should be developed and demonstrated using those records. After handoff, an authorized Unlock user will run the workflow against the real files. ACCESS MODEL At minimum we want: Access to Raw/Identity/Analysis layers for Unlock data custodian/admin. But we want access restricted to Analysis layer for Analyst + Future AI/Analysis Roles. We are not asking for Claude/Anthropic integration as part of this engagement. DELIVERABLES At the end of the engagement, we should have: 1. Secure Raw, Identity and Analysis storage within our existing AWS environment. 2. IAM policies/roles enforcing the access boundaries. 3. Appropriate S3 encryption/public-access/security settings. 4. Lightweight Python pseudonymization utility. 5. Stable random student-token/crosswalk functionality. 6. Configurable direct-identifier removal. 7. Basic QA output. 8. Synthetic test demonstrating the full workflow. 9. Short README explaining how to run the process. 10. Simple one-page architecture diagram. 11. 30-minute handoff/walkthrough. The workflow should ultimately be simple enough for us to run ourselves, e.g.: python pseudonymize.py --project sjc --input ./incoming EXPLICITLY OUT OF SCOPE We do not need a data lake, warehouse, database, VPC redesign, automated SharePoint ingestion, Airflow, Snowflake/Redshift, dashboards, web application, Claude/AI integration, hosted analytical environment, enterprise SSO, or full FERPA compliance audit. We are deliberately building a minimum viable secure workflow that can mature later. IDEAL BACKGROUND We're looking for someone senior enough to make good security decisions without overengineering the solution. Strong experience with: • AWS S3 • IAM / least-privilege access • AWS encryption/KMS • Python/pandas • secure data processing • pseudonymization/tokenization Experience with FERPA, HIPAA, education, healthcare or other sensitive-data environments would be particularly valuable. We value simple, secure and well-documented implementation over enterprise complexity.
Deschide pe Upwork