
Completed
Posted
Paid on delivery
Modify Databricks PySpark ETL Notebook with Dimension Joins and Aggregation Changes --- ## **Project Title** Modify Databricks PySpark ETL Notebook with Dimension Joins and Aggregation Changes --- ## **Project Description** I have an existing Databricks PySpark ETL notebook that creates a Delta table from a source fact table. I need to modify the transformation logic based on updated business requirements. ### Existing Environment * Databricks (Serverless) * PySpark * Unity Catalog * Delta Lake * Existing notebook uses metadata framework, DQ framework, AddAudits(), ValidateSchema(), DynamicOverwrite(), etc. --- ### New Requirements Need to enrich the fact data using three dimension tables. ### 1. Business Mapping Join: ``` fact.material_number = gold_catalog.commercial.dim_arkieva_product_line_cdb.material_cd ``` Bring: ``` business_column_j_as_per_combined_data ``` Rename it as ``` Business ``` This replaces the existing Business value. --- ### 2. Macro Ship To Mapping Join: ``` fact.ship_to_customer = gold_catalog.commercial.dim_ship_to_category_cdb.ship_to ``` Bring ``` macro_ship_to ``` --- ### 3. Macro Sold To Mapping Join ``` fact.customer_number = gold_catalog.commercial.dim_sold_to_category_cdb.sold_to ``` Bring ``` macro_sold_to ``` --- ### Important Business Change The output should **NOT** contain: ``` ship_to_code sold_to_code ``` Instead, the final aggregation must be rolled up using ``` Business Plant Plant Name Producer Location macro_ship_to macro_sold_to Material ``` instead of Ship To / Sold To codes. The aggregation logic therefore needs to be updated accordingly. --- ### Required Changes * Modify transformation logic * Add three dimension joins * Update aggregation/groupBy * Update final DataFrame projection * Update Delta table DDL * Update StructType schema * Update metadata/primary key if required * Ensure ValidateSchema() and EnforcePK() pass * Ensure notebook runs successfully on Databricks Serverless --- ### Deliverables * Updated transformation code * Updated schema * Updated DDL * Any required PK/metadata changes * Code should be clean, production-ready and well commented. ---
Project ID: 40625547
4 proposals
Remote project
Active 2 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs

With over 6 years of experience in Data Integration, Data Modeling and ETL using SQL, I have the expertise to tackle the complex task of modifying your existing Databricks PySpark ETL notebook to include dimension joins and aggregation changes. I am well-versed with Delta Lake, Unity Catalog and PySpark and have a solid understanding of the metadata framework and DQ framework your notebook utilizes. In line with your project requirements, I will ensure that the transformed fact data is enriched with the three dimension tables. By rewriting the transformation logic, adding dimension joins and updating aggregation/groupBy processes, I will guarantee that your output correctly rolls up using 'Business', 'Plant', 'Plant Name', 'Producer Location', 'macro_ship_to' and 'macro_sold_to'. Moreover, I will diligently update Delta table DDL, schema and metadata/primary key if required to ensure ValidateSchema() and EnforcePK() pass, enabling smooth execution on your Databricks Serverless environment. My commitment to clean, production-ready code paired with 100% on-time delivery rate make me an ideal choice for this project. Let's work together to enhance your existing ETL notebook to meet your updated business requirements. Looking forward to making a positive impact on your data transformation tasks!
₹4,000 INR in 7 days
1.4
1.4
4 freelancers are bidding on average ₹3,675 INR for this job

I can't tell you how excited I was to stumble upon this project, as it aligns so well with my expertise and experience. As a seasoned developer, I've immersed myself in different aspects of data management and SQL for years, giving me the extensive understanding needed to make your project a resounding success. Additionally, my proficiency in popular tools like Databricks PySpark, Unity Catalog, and Delta Lake ensures that I'm comfortable working within the existing environment. On the technical front, throughout my career I have developed a deep understanding of ETL processes and it's in my wheelhouse to join multiple dimension tables as you have requested. Updating the aggregation/logic accordingly is also second nature to me. Importantly, my pragmatic mindset will keep rollup logic easy but correct: using Business Plant PlantName ProducerLocation macroShipTo macroSoldTo Material instead of Ship To/Sold To codes. Lastly, reliability is at the core of what I deliver. This means clean, production-ready code with comprehensive comments so your team can take over the notebook with ease. My work doesn’t end at delivery—I stand by my projects and will continue supporting you as and when needed. So let’s work together to create an ETL notebook that not only ticks off your current needs but sets the ground for future modifications.
₹3,500 INR in 2 days
2.6
2.6

Hyderabad, India
Member since Sep 10, 2022
₹2000-3500 INR
₹600-1500 INR
₹600-5000 INR
₹600-1500 INR
₹600-1500 INR
₹100-400 INR / hour
$15-25 USD / hour
₹600-1500 INR
$15-25 USD / hour
₹600-1500 INR
₹750-1250 INR / hour
€30-250 EUR
$15-25 USD / hour
$30-250 USD
$250-750 USD
₹750-1250 INR / hour
$30-250 AUD
$30-250 USD
₹1500-12500 INR
$250-750 USD
$10-30 AUD
$750-1500 USD
$40-60 USD
$10-30 USD
$30-250 USD