Data Engineer
Current* Proficient in Apache Spark development using PySpark for scalable data processing and analysis.* Extensive development experience in Apache Spark using PySpark and Scala for large-scale data processing and analytics and developed and optimized Spark jobs for real-time and batch processing, enabling data-driven insights.* Extensive experience with Azure Databricks platform, leveraging its collaborative environment to build and optimize Spark workflows.* Designed and implemented Spark jobs for complex data transformations, data enrichment, and machine learning tasks.* Expertise in utilizing Azure Data Factory to orchestrate end-to-end data workflows.* Designed data pipelines to extract, transform, and load (ETL) data from diverse sources into Azure Databricks and other target destinations.* Leveraged Azure Data Factory's data movement activities to efficiently transfer data between Azure services.* Implemented data storage solutions using Azure Storage, including Blob storage and Data Lake Store.* Leveraged Databricks Delta for managing large-scale datasets, enabling ACID transactions, and optimizing data storage and querying performance.* Developed data archival strategies to optimize storage costs and ensure data availability as per business requirements.* Proficient in Git-based version control, maintaining clean and organized repositories for codebase management.* Exposure to Azure DevOps for setting up continuous integration and continuous deployment (CI/CD) pipelines.*Automated deployment processes using Azure DevOps, ensuring efficient and reliable application releases.