Data Engineer
Current• Work on requirements gathering, analysis and designing of the systems.• Developed Spark programs using Scala to compare the performance of Spark with Hive and SparkSQL.• Developed spark streaming application to consume JSON messages from Kafka and perform transformations.• Involved in monitoring Big query, Dataproc and cloud Data flow jobs via Stack driver for all the different environments.• Consulting with client business community to gather requirements for specific BI scorecard, vdashboard and reporting needs.• Wrote detailed specifications to be used in developing reports and processes. • Worked with product teams to create various store level metrics and supporting data pipelines written in GCP’s bigdata stack.• Used Sqoop import/export to ingest raw data into Cloud Storage by spinning up Cloud Dataproc cluster.• Developed dashboard prototypes using Cloud Dashboard Tools Looker and AWS Quicksight managing all aspects of the technical development.• Preparing design documents for changes to be implemented, performing peer reviews and developing scripts to load and process data. • Applying technical skills using Python, VBA, Tableau, Power BI, Essbase and SQL in data collection, data analysis and reporting to procure data from database structures and provide Business Intelligence solutions to client requests in a timely manner. • Worked on migrating an entire oracle database to BigQuery and using of power bi for reporting.• Worked on moving data between GCP and Azure using Azure Data Factory.• Created and maintain the data pipelines using Matillion ETL and Fivetran.• Created measure and dimensions on Looker based on the data and the business requirements.• Converted PL/SQL type of code to both bigquery - python architecture as well as azure databricks and pyspark in dataproc.