← Zakázky

Build a 1,500+ File Synthetic Microsoft 365 Data Estate

Rozpočet: $500.0 FIXED / ⭐ 4.87 (7) United States

data-mining, administrative-support, data-scraping

Budget: $500 fixed price Timeline: 2–4 days Environment: Microsoft 365 Developer Tenant and SharePoint Online Project Overview We need an experienced Microsoft 365 and synthetic-data resource to build a substantial demonstration data estate for Smart Stack 3.0. The environment will contain three distinct synthetic companies representing: Financial services Healthcare Manufacturing Each company should have its own realistic organizational structure, departments, SharePoint sites, document libraries, folders, users, and business records. The purpose is to demonstrate how the Smart Stack 3.0 Refinery can examine a complex enterprise data estate and surface valuable records, sensitive information, duplicates, obsolete content, inconsistent classifications, OCR requirements, and other data-quality conditions. This is not a request for 1,500 random files. The finished environment must feel like three believable organizations with connected records, recurring entities, document histories, and intentionally designed data-quality scenarios. All information must be entirely synthetic. No real personal, patient, financial, customer, employee, or confidential information may be used. Required Volume Create and upload a minimum of 1,500 files: Financial services: minimum 500 files Healthcare: minimum 500 files Manufacturing: minimum 500 files We expect the contractor to use automation, scripts, structured templates, and batch generation to achieve this volume within the fixed budget. The files should include a practical mix of: PDFs Scanned or image-based PDFs Excel spreadsheets CSV files Word documents PowerPoint presentations Images or supporting attachments Text and log files where appropriate No single file type should dominate the entire environment. Company and SharePoint Structure Create three clearly differentiated synthetic companies. Each company should include: A unique company name and basic profile Approximately 50–100 synthetic users Multiple departments and business functions At least three SharePoint sites Multiple document libraries and folder structures Realistic ownership, modification dates, and file naming Documents spanning multiple business years A mixture of well-managed and poorly managed locations Suggested departments include: Executive Finance Legal and compliance Human resources Operations Sales and customer service Information technology Procurement Risk or quality management Industry-specific business functions Financial Services Data Set Create realistic synthetic content such as: Customer and account documentation Loan and mortgage records Transaction and reconciliation reports Investment and portfolio reports Know Your Customer documentation Risk assessments Regulatory and compliance reviews Internal audits Vendor records Policies and procedures Board and executive reports Financial projections and operating reports Include fictional account numbers, taxpayer identifiers, payment details, customer information, and other sensitive-looking data that can be safely detected during testing. Healthcare Data Set Create realistic synthetic content such as: Patient registration records Clinical encounter summaries Claims and billing files Provider documentation Insurance and eligibility records Facility and operational reports Privacy and compliance reviews Policies and procedures Vendor agreements Staffing and scheduling records Quality-of-care reports Administrative spreadsheets Include synthetic health information, patient identifiers, insurance details, contact information, and other sensitive-looking content. Nothing may relate to a real person. Manufacturing Data Set Create realistic synthetic content such as: Bills of materials Product specifications Engineering documents Production schedules Quality-control records Equipment maintenance logs Inventory reports Supplier documentation Purchase orders and invoices Safety and incident reports Shipping and logistics records Plant operating reports Customer and warranty records Documents should share realistic relationships, such as matching product numbers, suppliers, facilities, purchase orders, and production runs. Required Refinery Test Scenarios The data estate must deliberately include: Exact duplicate files Near-duplicate files Multiple versions of the same record Draft, final, and superseded documents Redundant, obsolete, and trivial content High-value authoritative records Poor and inconsistent filenames Misfiled documents Missing or inconsistent metadata Documents containing conflicting information Sensitive-looking synthetic information Structured information embedded in unstructured documents Scanned PDFs requiring OCR Documents with signatures, stamps, or handwritten-looking elements Empty, corrupted, or unreadable test files Unusually large spreadsheets Files with old modification dates Records with different retention requirements Documents shared across departments Data that should and should not be prepared for AI use At least 30% of the environment should participate in a documented test scenario, rather than existing only as filler. Ground-Truth Requirements Provide a master inventory in Excel or CSV covering every generated file, including: Company and industry SharePoint site Library and folder Filename File type Department Document category Synthetic owner Created and modified dates Sensitivity category Expected retention category Business-value category Duplicate or version family ROT designation OCR requirement Authoritative-record designation Intended Smart Stack test scenario The ground-truth inventory is a critical project deliverable. It will allow us to compare Smart Stack 3.0 results against the expected conditions built into the data. Deliverables The selected contractor must provide: Three complete and differentiated synthetic company data sets A minimum of 1,500 total files At least three SharePoint sites per company All content uploaded and organized in our Microsoft 365 Developer environment A complete ground-truth inventory A summary of the major test scenarios All scripts, templates, configuration files, and generation instructions A short handoff guide explaining how to regenerate or expand the data Confirmation that the environment contains no real personal or confidential data Acceptance Criteria The project will be accepted when: At least 1,500 files have been created and uploaded Each industry contains at least 500 files The three companies are visibly and operationally distinct File content is readable and more than filenames with placeholder text Records use internally consistent fictional people, companies, dates, products, and identifiers Required duplicates, versions, sensitive-data examples, ROT content, and OCR scenarios are present At least 30% of files are tied to documented test scenarios The ground-truth inventory accurately matches the uploaded files All generation scripts and templates are transferred to us A scan or review confirms that no real personal or regulated data was used Ideal Candidate The ideal candidate has experience with: SharePoint Online and Microsoft 365 Developer tenants Synthetic enterprise data generation PowerShell or Python automation Microsoft Graph or SharePoint upload automation Word, Excel, PowerPoint, PDF, and CSV generation Information governance and records management Microsoft Purview concepts Financial services, healthcare, or manufacturing data Application Questions Please answer all five questions: How will you generate 1,500 realistic files within the $500 fixed budget? What tools or scripts will you use to create and upload the files? How will you make each industry data set internally consistent and clearly differentiated? How will you create and document duplicates, outdated records, sensitive data, OCR examples, and other test scenarios? Please provide an example of a similar Microsoft 365, SharePoint, or synthetic-data project. Budget Expectations This is a $500 fixed-price project. The work should be completed through intelligent automation and repeatable templates. We are not looking for 1,500 manually authored documents. We are looking for a well-designed, automated synthetic data estate that is large enough and realistic enough to demonstrate the value of the Smart Stack 3.0 Refinery.
Otevřít na Upwork