← Jobs

Extract WallStreetBets Data from Reddit Monthly Dumps

Budget: - HOURLY / PART_TIME ⭐ 4.81 (6) Chile

data-extraction

I'm looking for someone to help process the Reddit monthly dumps available on Academic Torrents: https://academictorrents.com/browse.php?search=Reddit Project The Reddit data is split into monthly comments (RC) and submissions (RS) files. Each monthly file contains data for all subreddits. I only need data from the r/wallstreetbets subreddit. The tasks are: Download all Reddit monthly dumps from 2024 onward (both submissions and comments). Extract only records where subreddit == "wallstreetbets". Discard all other subreddits. Produce one merged CSV (or equivalent) containing all WallStreetBets submissions. Produce one merged CSV containing all WallStreetBets comments. Ensure the output schema remains consistent across all months. Existing Code I already have a Python script that processes a single month. Your job is mainly to automate and extend it to all months from 2024 onward, merge the outputs, and verify that everything completed successfully. The existing script filters r/wallstreetbets and exports the desired fields. Deliverables Python script that processes all monthly files automatically. One merged submissions dataset (2024–present). One merged comments dataset (2024–present). Brief instructions explaining how to rerun the process for future months. Requirements Experience processing very large datasets (compressed .zst Reddit dumps). Python experience. Familiarity with streaming large files (preferred). Ability to verify that no months are missing. This is an academic research project, and unfortunately we have a limited budget, so please keep that in mind when submitting your quote. If you've done similar large-scale Reddit or data engineering work before, please mention it in your proposal. Thank you!
Open job