Data Engineer Needed to Build ETL/ELT Pipeline & Automate Data Warehouse Reporting
Rozpočet: $500.0
FIXED /
⭐ 5.00 (10)
Ireland
etl-pipelines, pyspark, apache-kafka, databricks-platform, snowflake, apache-spark, apache-airflow-platform, azure-devops
Preferované kvalifikace
- Zkušenost: Expert
We are looking for an experienced Data Engineer to design, build, and maintain a scalable ETL/ELT pipeline that extracts data from multiple sources, transforms it, and loads it into our cloud data warehouse for reporting and analytics.
Project Overview:
Our data currently lives in scattered sources (APIs, databases, CSV/Excel files, and third-party platforms) and needs to be consolidated into a single, reliable, automated pipeline. We need someone who can build a clean, well-documented data pipeline from raw data ingestion through to analytics-ready tables.
Responsibilities:
Design and implement ETL/ELT workflows to extract, transform, and load data
Build and optimize a data warehouse schema (star/snowflake schema, fact & dimension tables)
Set up data ingestion from REST APIs, SQL/NoSQL databases, and flat files (CSV, JSON, Excel)
Automate workflows using Apache Airflow (or Prefect/Dagster) for orchestration and scheduling
Write efficient, production-grade SQL and Python for data transformation
Implement data modeling best practices (dbt preferred) for clean, testable transformations
Ensure data quality, validation, and error-handling across the pipeline
Work with cloud platforms: AWS (S3, Redshift, Glue, Lambda), Google Cloud (BigQuery, Dataflow), or Azure (Data Factory, Synapse)
Optimize for performance and cost efficiency in data storage and query performance
Set up CI/CD for pipeline deployment and version control (Git)
Handle both batch processing and, if needed, real-time streaming using Kafka or Kinesis
Build or maintain a data lake / lakehouse architecture (Delta Lake, Databricks a plus)
Document data flow, schema, and pipeline architecture clearly for the team
Requirements:
Proven experience as a Data Engineer or similar role
Strong skills in SQL, Python, and pipeline orchestration tools (Airflow, dbt, Luigi)
Hands-on experience with cloud data platforms (AWS, GCP, or Azure)
Experience with data warehousing solutions (Snowflake, BigQuery, Redshift, or Synapse)
Understanding of data modeling, normalization, and schema design
Familiarity with big data tools (Spark, Hadoop) is a plus
Experience with API integrations and working with structured/unstructured data
Knowledge of data governance, security, and compliance best practices
Strong problem-solving skills and ability to work independently
Good communication skills and ability to document technical work clearly
Nice to Have:
Experience with Databricks or Snowflake
Familiarity with Terraform or Infrastructure-as-Code
Experience integrating with BI tools (Power BI, Tableau, Looker)
Background in machine learning pipeline support (MLOps) is a bonus
Deliverables:
Fully functional, automated ETL/ELT pipeline
Clean, queryable data warehouse tables ready for BI/reporting
Documentation of architecture, data flow, and maintenance instructions
Basic monitoring/alerting setup for pipeline failures
If you have a strong background in data engineering, pipeline automation, and cloud data infrastructure, we'd love to see examples of similar projects you've completed. Please share relevant experience with Airflow, SQL, Python, and cloud data warehouses in your proposal.
Otevřít na Upwork
AI proposal draft
Generate a short cover letter for this job. Edit before sending.
Sign in to generate an AI proposal draft.
Přihlásit