Data Engineer
Chicago, Illinois, Us
• Developed an autonomous Python ETL pipeline featuring unit testing, comprehensive error reporting, and a manual data intake process, feeding into GCP PostgreSQL• Engineered a relational PostgreSQL database using a structured, iterative design approach and successfully deployed it on Google Cloud Platform with a local/server persistent proxy authentication system• Leveraged CI/CD processes to engineer a cost-efficient Flask e-signature application and deployed a docker on Google Cloud Run. This strategic approach significantly reduced the tool's cost from $1,000 per user to a few cents.• Integrated various REST APIs to extract and/or load client data, including Box file cloud storage, close.io CRM, Google Sheets, and Gmail, into our Python-based data infrastructure.• Conducted extensive data conversions, transitioning unstructured data into standardized SQL formatsfor seamless integration into our Big Data platform. • Developed the department's automation architecture from scratch while also enhancing management flexibility with interface-to-code collaborative features.• Designed a repository structure that is logical, modular, and scalable, incorporating environmental files for secure credential management. This configuration significantly streamlined the process of applying repetitive code solutions, reducing the time investment from hours to just minutes.• Implemented an agile project management system and analyzed its insights for process enhancement. This led to greatly improved clarity and communication between the CEO and our team.• Maintained security and data governance options including GCP permissions, credentials storage, and sensitive data access through the FERPA security framework.• Addressed long standing data integrity issues which enabled capture of previously lost clients; adding 7% leads• Managed and processed unclean datasets up to 60 GB in size of types xlsx, csv, json, xml, and pdf.