Data Engineer
• Engineered and optimized data pipelines, resulting in a 30% improvement in data processing efficiency.• Implemented ETL processes using Python and Apache Airflow, reducing data latency by 25%.• Designed and maintained scalable data architectures on AWS, leading to a 20% reduction in storage costs.• Collaborated with cross-functional teams to integrate data from diverse sources, enhancing data accessibility by 40%.• Developed and managed data models in Snowflake, improving query performance by 50%.• Utilized SQL and NoSQL databases to support advanced analytics, increasing data retrieval speeds by 35%.Job Profile:• Developed and optimized data pipelines using Python, Java, and Scala to process large-scale datasets efficiently.• Deployed and managed data solutions on cloud platforms including AWS, Azure, and Google Cloud Platform (GCP) to ensure scalable and reliable data infrastructure.• Utilized Apache Spark, Hadoop, and Kafka for real-time data processing and analysis, enabling swift and informed decision- making.• Created automated workflows with Airflow and Autosys, integrating data from Snowflake, Redshift, and Teradata SQL Assistant to streamline data operations.• Implemented data storage and retrieval solutions using HDFS, Hive, Pig, and Sqoop within Hadoop distributions such as Hortonworks and Cloudera Hadoop.• Leveraged Spark, Spark SQL, and Kafka for efficient real-time data processing and analytics across Linux, Unix, ZOS, and Windows operating systems.• Designed and executed ETL workflows with IBM Infosphere Information Server, ensuring seamless data integration and transformation across various platforms.