Lead Data Engineer
Current• Build complex ETL jobs that transform data visually with data flows or by using compute services Azure Databricks, and Azure SQL Database• Develop and maintain various data ingestion pipelines as per the design architecture and processes: source to landing, landing to curated & curated to process • Use various types of activities: data movement activities, transformations, and control activities; Copy data, Data flow, Get Metadata, Lookup, Stored procedure, Execute Pipeline• Dealt with Storage services named Data Lake Storage Gen 1 and Gen 2, Storage Explorer for hosting CSV, JSON and Parquet files and managing access across storage accounts.• Write Databricks notebooks (Python) for handling large volumes of data, transformations, and computations• Utilize Databricks Delta Lake storage layer to create versioned Apache Parquet (delta) files with transaction log and audit history• Work with various file formats: flat-file TXT & CSV; parquet & other compressed formats• Build Delta Lake for the curated layer, maintain high-quality data available for the teams: data scientists, finance etc.• Utilize Azure’s ETL, Azure Data Factory (ADF) services to ingest data from legacy disparate data stores - SAP (Hana), SFTP servers & Cloudera Hadoop’s HDFS to Azure Data Lake Storage (Gen2)• Automated dataflows using Logic apps and Power Automate (Flow) which connects different Azure services and Function apps for customizations.• Scheduled and maintained Azure Databricks jobs.