I will build etl pipelines using databricks, snowflake and pyspark
Data Engineer, ETL Pipelines with Databricks and Snowflake
Over deze dienst
Struggling with messy data scattered across multiple sources?
I build reliable, production-ready ETL pipelines using Databricks, Snowflake, and PySpark turning raw, disconnected data into clean, structured, analytics-ready warehouses.
With 5 years of hands-on data engineering experience, I help businesses:
Extract data from multiple sources (APIs, databases, files, cloud storage)
Transform and clean data using PySpark for accuracy and performance
Load data into Databricks, Snowflake, or Microsoft Fabric
Automate and orchestrate pipelines for scheduled, hands-off runs
Optimize existing pipelines for speed and cost efficiency
Whether you need a simple single-source pipeline or a full end-to-end data architecture, I deliver clean, well-documented code and clear communication throughout.
Why work with me:
- 5 years professional data engineering experience
- Expertise in Databricks, Snowflake, Fabric, PySpark, SQL
- On-time delivery with detailed documentation
- Fast, clear communication
Let's build a pipeline that gets your data working for you. Message me before ordering if you have a custom requirement happy to scope it out together.
Veelgestelde vragen
Q: What tools and platforms do you work with?
I specialize in Databricks, Snowflake, Microsoft Fabric, PySpark, and SQL. I also work with Azure Data Factory, dbt, and Apache Airflow for orchestration when needed.
Q: What if I need changes after delivery?
A: Each package includes revisions (see package details above). I also offer ongoing support and maintenance as a custom add-on if you need it.
Q: Do you work with large datasets / big data?
A: Yes, PySpark and Databricks are built for large-scale, distributed data processing, so I can handle datasets that don't fit in traditional tools like Excel or basic SQL.
Q: I'm not sure which package I need. Can we discuss my project first?
A: Absolutely, message me with details about your data sources, destination, and goals, and I'll recommend the right package or a custom offer.
Q: Can you automate the pipeline to run on a schedule?
A: Yes, I can set up orchestration so your pipeline runs automatically on a schedule (daily, hourly, etc.) without manual intervention — available in the Standard and Premium packages.
Q: Can you build a pipeline that pulls from multiple data sources?
Yes. I build ETL pipelines that combine data from APIs, databases, cloud storage, and flat files into a single, clean warehouse in Snowflake or Databricks.
