Data Engineer
CurrentConstructed a data pipeline that ingest data to HDFS from multiple sources. Led the performance initiatives that optimizes the queries to yield a 40% improvement in data processing speed. Rather than analyzing data in a database, move this data to a data lake and analyze it. Optimized data storage solutions, resulting in a 25% reduction in storage cost without data integrity. Wrote a bunch of transformation in one of the existing pyspark project.