Cloud Data Engineer
Current•Performed data investigation to discover correlations trends and the ability to explain them. Worked with Data Engineers, Data Architects, to define back-end requirements for data products (aggregations, materialized views, tables – visualization) •Performing data analysis, statistical analysis, generated reports, listings, and graphs using SAS tools, SAS/Graph, SAS/SQL, SAS/Connect and SAS/Access. •Developing Spark applications using Scala and Spark-SQL for data extraction, transformation, and aggregation from multiple file formats. Using Kafka and integrating with the Spark Streaming. •Migrate data from on-premises to AWS storage buckets. Developed data analysis tools using SQL and Python code. Developed various Mappings with the collection of all Sources, Targets, and Transformations using Informatica Designer. •Designed and implemented Sqoop for the incremental job to read data from DB2 and load to hive tables and connected to Tableau for generating interactive reports using Hive server2. •Used Spark Streaming to receive real time data from the Kafka and store the stream data to HDFS using Python and NoSQL databases such as HBase and Cassandra. •Collected data using Spark Streaming from AWS S3 bucket in near-real-time and performs necessary Transformations and Aggregation on the fly to build the common learner data model and persists the data in HDFS. Used Apache NiFi to copy data from local file system to HDP. •Worked on Dimensional and Relational Data Modeling using Star and Snowflake Schemas, OLTP/OLAP system, Conceptual, Logical and Physical data modeling using Erwin. •Automated the data processing with Oozie to automate data loading into the Hadoop Distributed File System. •Experience in Converting existing AWS Infrastructure to Server less architecture (AWS Lambda, Kinesis), deploying via Terraform and AWS Cloud Formation templates. •Architect and design server less application CI/CD by using AWS Server less (Lambda) application model.