Senior AI/LLM Dataset Engineer
Бюджет: $15.0 - $30.0
HOURLY / PART_TIME
⭐ 4.90 (5)
USA
numpy, data-science, python, machine-learning, artificial-intelligence, data-analysis, python-script
Preferred qualifications
- Location: India, United States
- Experience: Expert
Role Overview:
The position focuses on building and validating golden datasets for an AI/LLM-based Natural Language Query (NLQ) Evaluation Framework. You'll be responsible for designing synthetic data generation pipelines, developing SQL-based QA templates, validating datasets, and creating schema documentation while collaborating with domain experts and engineering teams.
Key Responsibilities:
Design and implement synthetic data generation pipelines using Python, NumPy, Faker, and pandas.
Generate realistic datasets with statistical distributions and controlled data imperfections.
Validate datasets using DuckDB and maintain dataset versioning.
Develop Jinja2 SQL templates and verify expected answers against reference SQL.
Collaborate with Domain SMEs to create QA datasets and support LLM evaluation.
Create ER diagrams, DDL scripts, Data Dictionaries, and CSV Header Specifications.
Required Skills:
Strong Python programming with NumPy, pandas, and SQL
Experience with synthetic data generation and data validation
Knowledge of DuckDB, Jinja2, and Git
Understanding of database design, schema documentation, and SQL development
Exposure to AI/LLM evaluation frameworks is an added advantage
Отвори в Upwork
AI proposal draft
Generate a short cover letter for this job. Edit before sending.
Sign in to generate an AI proposal draft.
Вход