Senior Data Engineer / Lead Data Engineer — Azure Data, Fabric & Power BI Platform
Buget: -
HOURLY / FULL_TIME
⭐ 4.76 (3)
United States
data-migration, data-source-integration, windows-azure, sql, microsoft-power-bi, etl-pipelines, python, sql-azure
Preferred qualifications
- Experience: Intermediate
Senior Data Engineer — Microsoft Azure, Fabric & Power BI (Offshore)
Job Title: Senior Data Engineer / Lead Data Engineer — Azure Data Platform Department: Data & Analytics / Technology Reports To: Director of Data Engineering / Head of Data & Analytics Location: Offshore (remote) — [India / Philippines / LATAM / specify] Engagement Type: [Full-time employee / Contract / Contract-to-hire] — [duration- 6-9 months (initially)] Working Hours: Minimum 4–5 hours of daily overlap with U.S. Eastern Time (approx. 8:00 AM – 1:00 PM ET); flexibility required during migration cutovers, month-end close, and quarter-end reporting cycles Experience Required: 6–8+ years Compensation: [Rate / band (Open for discussion)]
________________________________________
Position Summary
We are seeking a hands-on Senior Data Engineer to design, build, and operate a modern cloud data platform on Microsoft Azure and Microsoft Fabric. This person will own the full data lifecycle — from source system extraction and migration through transformation, orchestration, modeling, and delivery into Power BI semantic models consumed by finance, asset management, investments, and executive leadership.
The ideal candidate is equally comfortable architecting a medallion lakehouse, debugging a 300-line stored procedure, tuning a slow DAX measure, and explaining a data lineage issue to a non-technical accountant. They should be able to work with a high degree of autonomy across time zones, translate ambiguous business requirements into durable data assets, and bring opinions about how a data platform should be built rather than waiting for a specification.
This is an offshore role embedded in a distributed team. Strong written English, proactive communication, and disciplined documentation are as important as technical depth.
________________________________________
Key Responsibilities
Data Architecture & Platform Engineering
• Design and evolve the enterprise data architecture on Azure — lakehouse / medallion (Bronze–Silver–Gold) patterns, dimensional models, data vault where appropriate, and semantic layers.
• Build and maintain data assets across Azure SQL Database, Azure Data Lake Storage Gen2, Microsoft Fabric (OneLake, Lakehouse, Warehouse, Dataflows Gen2, Direct Lake), Azure Synapse Analytics, and Azure Databricks.
• Make and defend platform decisions — Fabric vs. Synapse vs. Databricks for a given workload, Delta vs. Parquet, warehouse vs. lakehouse endpoint, import vs. Direct Lake vs. DirectQuery.
• Define and enforce standards for naming, layering, partitioning, schema evolution, and environment promotion.
Data Migration & Modernization
• Lead end-to-end migrations: on-premises SQL Server / legacy warehouses / Access / flat-file estates → Azure SQL, Synapse, Fabric, or Databricks.
• Perform source system discovery, data profiling, gap analysis, field-level mapping, and migration runbook development.
• Build reconciliation and validation frameworks (row counts, control totals, financial tie-outs, hash comparisons) to prove migrated data matches source-of-truth.
• Plan and execute cutovers with rollback strategies, parallel-run periods, and stakeholder sign-off gates.
Data Integration, Transformation & Orchestration
• Build scalable ingestion pipelines in Azure Data Factory / Fabric Data Factory using parameterized, metadata-driven patterns — not one-off copy activities per table.
• Ingest from REST APIs, SFTP, ODBC/JDBC, flat files, Excel workbooks, SharePoint, and vendor ERP exports; handle full loads, incremental/CDC/watermark loads, and late-arriving data.
• Develop transformation logic in PySpark / Spark SQL (Databricks notebooks), T-SQL stored procedures, and Fabric notebooks; implement Delta Lake features (MERGE/upsert, time travel, OPTIMIZE, VACUUM, Z-order/liquid clustering).
• Design and operate orchestration and workflow layers — ADF pipelines and triggers, Fabric Data Pipelines, Databricks Workflows/Jobs, dependency chains, retries, alerting, and SLA monitoring.
• Implement robust error handling, logging, audit tables, data quality checks, and automated failure notifications.
SQL Development & Performance
• Write advanced T-SQL: complex joins, window functions, CTEs, recursive queries, dynamic SQL, stored procedures, functions, temporal tables, and MERGE logic.
• Diagnose and remediate performance issues — execution plans, indexing strategy, statistics, partitioning, parameter sniffing, query store analysis, Spark shuffle/skew, and file compaction.
• Own cost and capacity management: Fabric capacity units (CU) consumption, Databricks cluster sizing and autoscaling policies, Synapse DWU management, storage tiering.
Semantic Modeling, DAX & Power BI
• Build enterprise-grade Power BI semantic models — star schemas, role-playing dimensions, calculation groups, aggregations, incremental refresh, composite models, and Direct Lake configurations.
• Author complex DAX: time intelligence, period-over-period and inception-to-date measures, dynamic segmentation, CALCULATE/filter-context manipulation, SUMMARIZECOLUMNS, and variables for readability and performance.
• Optimize model size and query performance using DAX Studio, Tabular Editor, VertiPaq Analyzer, and Performance Analyzer.
• Implement Row-Level Security (RLS) / Object-Level Security (OLS), workspace governance, deployment pipelines, and refresh scheduling.
• Partner with BI developers and business users on report design, but this role's center of gravity is the model and the data behind it.
Governance, Security & DevOps
• Implement data governance and lineage using Microsoft Purview, Fabric domains/endorsements, and/or Databricks Unity Catalog.
• Apply security best practices: Entra ID (Azure AD) authentication, managed identities, Key Vault secrets, private endpoints, RBAC, data masking, and PII handling.
• Manage source control and CI/CD via Git (Azure DevOps / GitHub) — branching strategy, pull request review, automated deployment of pipelines, notebooks, SQL objects, and Power BI artifacts across Dev/Test/Prod.
• Infrastructure-as-code exposure (Bicep, ARM, or Terraform) is expected at a working level.
Collaboration & Delivery
• Work directly with accounting, fund reporting, asset management, and investment teams to translate business logic into data models.
• Produce and maintain documentation: architecture diagrams, source-to-target mappings, data dictionaries, lineage maps, and runbooks.
• Participate in agile ceremonies, provide accurate estimates, flag risks early, and give clear written status updates given limited real-time overlap.
• Mentor junior engineers and enforce code review standards.
________________________________________
Required Qualifications
Area Requirement
Experience 6–8+ years in data engineering, with demonstrable ownership of data architecture, data migration, data transformation, and data orchestration initiatives (not just ticket-level pipeline maintenance)
Azure Core Azure SQL Database, Azure Data Lake Storage Gen2, Azure Data Factory, Azure Synapse Analytics — production experience, multi-year
Fabric Hands-on Microsoft Fabric: OneLake, Lakehouse, Warehouse, Data Pipelines, Dataflows Gen2, Direct Lake, capacity management
Databricks Spark/PySpark, Delta Lake, Databricks Workflows, cluster configuration and optimization; Unity Catalog a plus
SQL Expert-level T-SQL; strong data modeling (Kimball dimensional modeling, SCD Type 1/2, fact/dim design)
Power BI & DAX Advanced DAX authoring and optimization; enterprise semantic model design; RLS; Tabular Editor / DAX Studio
Programming Python (PySpark, pandas, requests) for data engineering; strong scripting fundamentals
DevOps Git-based version control, CI/CD for data assets, environment promotion discipline
Education Bachelor's degree in Computer Science, Engineering, Information Systems, or equivalent practical experience
Communication Fluent written and spoken English; able to run a requirements conversation with a U.S.-based finance stakeholder unassisted
Certifications (preferred, not required)
• Microsoft Certified: Fabric Analytics Engineer Associate (DP-600)
• Microsoft Certified: Azure Data Engineer Associate (DP-203)
• Databricks Certified Data Engineer Professional
• Microsoft Certified: Power BI Data Analyst Associate (PL-300)
________________________________________
Nice to Have — ERP & Industry Source Systems
Prior experience extracting from, integrating with, or modeling data out of any of the following is a meaningful advantage. We do not expect all of them — depth in two or three is more valuable than passing familiarity with all:
• Yardi (Voyager, Elevate, Investment Manager) — property management, GL, and rent roll data
• Investran (FIS) — fund accounting, partnership capital, waterfall and allocation data
• SS&C Precision LP — private equity / fund partnership accounting
• SS&C Advent Geneva — portfolio accounting and investment book of record
• HIA (Hotel Investor Apps) — hospitality-specific ERP and financial reporting
• Dynamo Software — investor CRM, fundraising pipeline, and investor reporting
• MRI Software — commercial property management and accounting
• Adjacent systems: Entrata, RealPage, Nexus/AvidXchange, Sage Intacct, NetSuite, Workday Adaptive, Anaplan, Salesforce, Opera PMS, STR/CoStar/Trepp data feeds
Practical experience with vendor data quirks — undocumented schemas, read-replica access, API rate limits, nightly export files, inconsistent GL segment structures, and multi-entity/multi-property consolidations — is highly valued.
________________________________________
Nice to Have — Domain Knowledge
Working knowledge of any of the following U.S. real estate and investment domains is advantageous:
• U.S. Commercial Real Estate: property/entity/fund hierarchies, rent rolls, leases and escalations, CAM reconciliations, NOI and cap-rate mechanics, budget vs. actual reporting, acquisitions and dispositions, JV structures and promote/waterfall calculations.
• Debt / Credit: whole loans, bridge and construction lending, mezzanine debt, preferred equity, CPACE, loan tape and servicing data, draw schedules, interest accruals and reserves, covenants, DSCR/LTV/debt yield metrics, maturity and extension tracking, warehouse lines and securitization reporting.
• Hotel / Hospitality Property Types: USALI-compliant P&L structures, ADR / Occupancy / RevPAR / RevPAR Index, STR competitive set reporting, brand and franchise data feeds, PIP and capex tracking, F&B and departmental reporting, management agreements.
Candidates without direct domain background but with strong financial data literacy and a demonstrated ability to learn a business quickly will still be considered.
________________________________________
Core Competencies
• Ownership mindset — sees a data problem through to resolution rather than handing it back to the business
• Precision with financial data — understands that a reconciliation break is not a rounding issue
• Clear asynchronous communication — writes well, documents proactively, escalates early
• Pragmatism — chooses the maintainable solution over the clever one
• Curiosity about the business — asks why a metric is calculated the way it is
________________________________________
Offshore Engagement Requirements
• Reliable high-speed internet with a stable backup connection and a quiet, professional work environment
• Willingness to work within the specified U.S. overlap window and to support periodic evening/weekend deployment or close-cycle activities
• Compliance with company information security policy: MFA, company-managed device or approved VDI access, VPN, no local storage of production data, and adherence to data residency and confidentiality requirements
• Ability to sign NDA and, where applicable, complete background verification
• Comfortable operating within a distributed team using Microsoft Teams, Azure DevOps/Jira, and shared documentation platforms
________________________________________
Success Measures
First 30 days — Environment access established; source systems, existing pipelines, and semantic models documented; first production pipeline fix or enhancement deployed.
First 90 days — Independently owning a defined data domain end-to-end; delivered at least one migration or new ingestion workstream; measurable improvement in refresh reliability or pipeline runtime.
First 6–12 months — Recognized as the technical owner of a major platform area; reduced manual/Excel-based reporting dependencies; established reusable, metadata-driven patterns adopted by the wider team.
________________________________________
Interview Process (suggested)
1. Recruiter screen — experience, availability, overlap hours, engagement terms
2. Technical screen — Azure/Fabric architecture discussion, SQL and DAX problem-solving
3. Practical exercise — small migration or transformation scenario with reconciliation logic
4. Panel — architecture deep dive plus a business-stakeholder communication conversation
5. Final — leadership fit and offer
Deschide pe Upwork
AI proposal draft
Generate a short cover letter for this job. Edit before sending.
Sign in to generate an AI proposal draft.
Autentificare