It Engineer Ii
Current• Spearheaded the complete design and implementation of the company's data lake house migration from HDFS to a cloud-native solution, leveraging MinIO for scalable object storage, and Apache Iceberg for efficient data querying. Developed robust ETL pipelines using PySpark for seamless data processing and integrated Airflow for automated workflow orchestration, driving operational efficiency and enabling advanced data analytics capabilities on a Kubernetes-based architecture.• Built data pipelines using Apache Spark to process large amounts of structured and unstructured data.• Developed automated processes for collecting, organizing, and analyzing big data.• Analyzed large datasets to identify patterns and trends in data.• Developed ETL jobs to extract, transform, and load data from various sources into the target system.• Provided technical support to business users on using Big Data tools for their analytical needs.• Collaborated with other members of the team to design efficient solutions for managing Big Data workloads.• Conducted research on emerging technologies related to Big Data processing, storage, and analytics.• Monitored production systems for identifying potential issues with performance or scalability.• Identified areas where automation can be used to streamline development workflows.• Implemented the gold, silver, and bronze layered architecture to structure data storage and processing, enhancing data governance and accessibility.• Cleaned and manipulated raw data.• Managed orchestration and workflow automation using Airflow, optimizing data processing pipelines, and improving operational efficiency.