Data Engineer
Current• Medallion Architecture (Bronze, Silver, Gold) using Azure Databricks, PySpark, and Delta Tables.Improve data quality, make the system more scalable, and ensured better data governance.• Built batch processing pipelines for time-series data, loading them into Delta Tables. • Implemented UAT and Production environments, ensuring seamless deployment and testing• Set up CI/CD pipelines with Azure DevOps, simplifying environment management, quality checks, and improving team collaboration.• Extracted data from various sources like CSV, Excel, PDFs, and JSONs (including nested ones) to create flat files used for model building.• Automated email alerts for data changes in Delta Tables using Logic Apps, Databricks, and API connectors.• Boosted notebook performance by optimizing code (runtime reduced from hours to minutes) and moved workloads to job clusters, which significantly cut down on costs.• Migrated and replaced SQL Server with Delta Tables, enhancing data governance and lineage• Automated several manual reports using Python (Pandas, NumPy), SharePoint, Data Lake, Logic Apps, and Databricks, tying it all together with workflow job triggers.• Managed and optimized Databricks clusters and Azure services, reducing monthly costs by ₹1.5 lakhs while keeping everything running smoothly and efficiently.