← Jobs

Senior AI/LLM Dataset Engineer

Budget: $15.0 - $30.0 HOURLY / PART_TIME ⭐ 4.90 (5) USA

numpy, data-science, python, machine-learning, artificial-intelligence, data-analysis, python-script

Preferred qualifications

  • Location: India, United States
  • Experience: Expert
Role Overview: The position focuses on building and validating golden datasets for an AI/LLM-based Natural Language Query (NLQ) Evaluation Framework. You'll be responsible for designing synthetic data generation pipelines, developing SQL-based QA templates, validating datasets, and creating schema documentation while collaborating with domain experts and engineering teams. Key Responsibilities: Design and implement synthetic data generation pipelines using Python, NumPy, Faker, and pandas. Generate realistic datasets with statistical distributions and controlled data imperfections. Validate datasets using DuckDB and maintain dataset versioning. Develop Jinja2 SQL templates and verify expected answers against reference SQL. Collaborate with Domain SMEs to create QA datasets and support LLM evaluation. Create ER diagrams, DDL scripts, Data Dictionaries, and CSV Header Specifications. Required Skills: Strong Python programming with NumPy, pandas, and SQL Experience with synthetic data generation and data validation Knowledge of DuckDB, Jinja2, and Git Understanding of database design, schema documentation, and SQL development Exposure to AI/LLM evaluation frameworks is an added advantage
Open job

AI proposal draft

Generate a short cover letter to copy into the offer. Says you are interested and ready to work.

Sign in to generate an AI proposal draft.

Log in