Senior Data Engineer
CurrentAs a Data Engineer, I’m applying pre-existing machine learning and data mining techniques to create predictive models. Developed high speed data ingestion pipeline using Scala, Apache Spark, Akka Streams, HDFS, Hive, and Cassandra. Developed python scripts for data cleaning, analysis and automating day to day activities. Hydrate the data lake by creating data ingestion pipelines for using various technologies likeAttunity Replicate, Kafka, Spark, HDFS, Hive, Java, Bash Scripting HBase, Teradata, Query Grid, CA7.The data is then used by stakeholders for machine learning and reporting. Used pandas, NumPy, Seaborn, matplotlib, Scikit-learn, SciPy, NLTK in Python for developing various machine learning algorithms. Setup storage and data analysis tools in Amazon Web Services (AWS) cloud computing infrastructure.