Data Engineer
• Regularly maintained the ETL pipeline scraping data from various multinational retail corporations for 20+ clients and transforming it into Alloy’s data model stored in BigQuery using Python.• Worked closely with the client-solutions team to design, build and/or reimplement 30+ scrapers and extraction models for the ETL pipeline using Python’s Selenium, Requests, Pandas, and built-in libraries.• Identified and implemented enhancements to the data pipeline, from more accurate detection of erroneous files based on file size, format, content, and encoding to optimized usage of Google’s geocoding API using Python and Java. • Increased the traceability of scraped files by automating the assignment of data sources to each file using Python’s SQLalchemy, and PostgreSQL.• Experienced working with SFTP, Electronic Data Interchange (EDI), React, Google Cloud Platform (GCP), and Jenkins.