Build a 1,500+ File Synthetic Microsoft 365 Data Estate
Rozpočet: $500.0
FIXED /
⭐ 4.87 (7)
United States
data-mining, administrative-support, data-scraping
Budget: $500 fixed price
Timeline: 2–4 days
Environment: Microsoft 365 Developer Tenant and SharePoint Online
Project Overview
We need an experienced Microsoft 365 and synthetic-data resource to build a substantial demonstration data estate for Smart Stack 3.0.
The environment will contain three distinct synthetic companies representing:
Financial services
Healthcare
Manufacturing
Each company should have its own realistic organizational structure, departments, SharePoint sites, document libraries, folders, users, and business records.
The purpose is to demonstrate how the Smart Stack 3.0 Refinery can examine a complex enterprise data estate and surface valuable records, sensitive information, duplicates, obsolete content, inconsistent classifications, OCR requirements, and other data-quality conditions.
This is not a request for 1,500 random files. The finished environment must feel like three believable organizations with connected records, recurring entities, document histories, and intentionally designed data-quality scenarios.
All information must be entirely synthetic. No real personal, patient, financial, customer, employee, or confidential information may be used.
Required Volume
Create and upload a minimum of 1,500 files:
Financial services: minimum 500 files
Healthcare: minimum 500 files
Manufacturing: minimum 500 files
We expect the contractor to use automation, scripts, structured templates, and batch generation to achieve this volume within the fixed budget.
The files should include a practical mix of:
PDFs
Scanned or image-based PDFs
Excel spreadsheets
CSV files
Word documents
PowerPoint presentations
Images or supporting attachments
Text and log files where appropriate
No single file type should dominate the entire environment.
Company and SharePoint Structure
Create three clearly differentiated synthetic companies. Each company should include:
A unique company name and basic profile
Approximately 50–100 synthetic users
Multiple departments and business functions
At least three SharePoint sites
Multiple document libraries and folder structures
Realistic ownership, modification dates, and file naming
Documents spanning multiple business years
A mixture of well-managed and poorly managed locations
Suggested departments include:
Executive
Finance
Legal and compliance
Human resources
Operations
Sales and customer service
Information technology
Procurement
Risk or quality management
Industry-specific business functions
Financial Services Data Set
Create realistic synthetic content such as:
Customer and account documentation
Loan and mortgage records
Transaction and reconciliation reports
Investment and portfolio reports
Know Your Customer documentation
Risk assessments
Regulatory and compliance reviews
Internal audits
Vendor records
Policies and procedures
Board and executive reports
Financial projections and operating reports
Include fictional account numbers, taxpayer identifiers, payment details, customer information, and other sensitive-looking data that can be safely detected during testing.
Healthcare Data Set
Create realistic synthetic content such as:
Patient registration records
Clinical encounter summaries
Claims and billing files
Provider documentation
Insurance and eligibility records
Facility and operational reports
Privacy and compliance reviews
Policies and procedures
Vendor agreements
Staffing and scheduling records
Quality-of-care reports
Administrative spreadsheets
Include synthetic health information, patient identifiers, insurance details, contact information, and other sensitive-looking content. Nothing may relate to a real person.
Manufacturing Data Set
Create realistic synthetic content such as:
Bills of materials
Product specifications
Engineering documents
Production schedules
Quality-control records
Equipment maintenance logs
Inventory reports
Supplier documentation
Purchase orders and invoices
Safety and incident reports
Shipping and logistics records
Plant operating reports
Customer and warranty records
Documents should share realistic relationships, such as matching product numbers, suppliers, facilities, purchase orders, and production runs.
Required Refinery Test Scenarios
The data estate must deliberately include:
Exact duplicate files
Near-duplicate files
Multiple versions of the same record
Draft, final, and superseded documents
Redundant, obsolete, and trivial content
High-value authoritative records
Poor and inconsistent filenames
Misfiled documents
Missing or inconsistent metadata
Documents containing conflicting information
Sensitive-looking synthetic information
Structured information embedded in unstructured documents
Scanned PDFs requiring OCR
Documents with signatures, stamps, or handwritten-looking elements
Empty, corrupted, or unreadable test files
Unusually large spreadsheets
Files with old modification dates
Records with different retention requirements
Documents shared across departments
Data that should and should not be prepared for AI use
At least 30% of the environment should participate in a documented test scenario, rather than existing only as filler.
Ground-Truth Requirements
Provide a master inventory in Excel or CSV covering every generated file, including:
Company and industry
SharePoint site
Library and folder
Filename
File type
Department
Document category
Synthetic owner
Created and modified dates
Sensitivity category
Expected retention category
Business-value category
Duplicate or version family
ROT designation
OCR requirement
Authoritative-record designation
Intended Smart Stack test scenario
The ground-truth inventory is a critical project deliverable. It will allow us to compare Smart Stack 3.0 results against the expected conditions built into the data.
Deliverables
The selected contractor must provide:
Three complete and differentiated synthetic company data sets
A minimum of 1,500 total files
At least three SharePoint sites per company
All content uploaded and organized in our Microsoft 365 Developer environment
A complete ground-truth inventory
A summary of the major test scenarios
All scripts, templates, configuration files, and generation instructions
A short handoff guide explaining how to regenerate or expand the data
Confirmation that the environment contains no real personal or confidential data
Acceptance Criteria
The project will be accepted when:
At least 1,500 files have been created and uploaded
Each industry contains at least 500 files
The three companies are visibly and operationally distinct
File content is readable and more than filenames with placeholder text
Records use internally consistent fictional people, companies, dates, products, and identifiers
Required duplicates, versions, sensitive-data examples, ROT content, and OCR scenarios are present
At least 30% of files are tied to documented test scenarios
The ground-truth inventory accurately matches the uploaded files
All generation scripts and templates are transferred to us
A scan or review confirms that no real personal or regulated data was used
Ideal Candidate
The ideal candidate has experience with:
SharePoint Online and Microsoft 365 Developer tenants
Synthetic enterprise data generation
PowerShell or Python automation
Microsoft Graph or SharePoint upload automation
Word, Excel, PowerPoint, PDF, and CSV generation
Information governance and records management
Microsoft Purview concepts
Financial services, healthcare, or manufacturing data
Application Questions
Please answer all five questions:
How will you generate 1,500 realistic files within the $500 fixed budget?
What tools or scripts will you use to create and upload the files?
How will you make each industry data set internally consistent and clearly differentiated?
How will you create and document duplicates, outdated records, sensitive data, OCR examples, and other test scenarios?
Please provide an example of a similar Microsoft 365, SharePoint, or synthetic-data project.
Budget Expectations
This is a $500 fixed-price project. The work should be completed through intelligent automation and repeatable templates.
We are not looking for 1,500 manually authored documents. We are looking for a well-designed, automated synthetic data estate that is large enough and realistic enough to demonstrate the value of the Smart Stack 3.0 Refinery.
Otvoriť na Upwork