Senior Data Engineer
CurrentI led the development of a complex data pipeline project, orchestrating various data operations and transformations. This pipeline encompassed actions like file movement, Sqoop data extraction from Teradata and SQL sources, and data export into Hive staging tables. I collaborated on ETL tasks, ensuring data integrity and pipeline stability, with a focus on data retrieval from file systems using Spark commands.Python was central to our project, with code written for exploratory data analysis using machine learning libraries like Scikit-learn, NumPy, Pandas, and Matplotlib. We implemented Snowflake's COPY and INSERT statements, integrating Snowpipe for real-time data ingestion and transforming raw data into actionable insights.Our team also leveraged Cloud services for data ingestion and transformation and building a pipeline. We employed Python scripts for data cleansing, mapping, aggregation, and quality reporting, facilitating effective data management.Throughout the project, I utilized a diverse technology stack, including Python, PySpark, Docker, Kafka, and various data formats, to deliver a robust and efficient data pipeline solution.