Data Engineer
Demonstrated expertise with advanced big data technologies including Spark, Hadoop, Spark SQL, Azure Data Factory, Databricks, AWS Glue, Athena, and AWS Lambda.Proficient in Data Warehousing and Data Lake concepts, with hands-on experience in Azure Data Lake Storage, AWS S3, and HDFS.Skilled in using Git for version control, Intellij for development, and PowerShell for scripting.Thorough understanding and application of DevOps principles, including CI/CD pipelines, releases, and proficient in branching and merging strategies.Analyze user requirements comprehensively, design, and proficiently develop ETL processes to efficiently load enterprise data into Data Warehouses.Develop, rigorously test, schedule, and meticulously orchestrate ETL processes using both ETL pipelines and PySpark code.Proficiently analyze data by querying it using PySpark and SQL to ensure impeccable data quality and promptly identify any discrepancies.Proficiently manage operational activities including monitoring daily runs across all environments (Test, Acceptance, Production), meticulously debug failures, and promptly rectify errors.Implement rigorous data governance practices, ensuring impeccable data quality and integrity, and meticulously maintaining data security and privacy standards.Collaborate closely with data scientists and analysts to deeply comprehend their data requirements, ensuring unfailing availability, reliability, and high performance of data systems.Possess real-time expertise and a robust understanding of mathematics, statistical analysis, and adept at both supervised and unsupervised learning, proficient in building and meticulously evaluating predictive models.