Hadoop Developer
Current• Responsible for building scalable distributed data solutions using Hadoop• Imported data using Sqoop to load data from MySQL to HDFS on regular basis from various sources.• Written multiple MapReduce programs to power data for extraction, transformation and aggregation from multiple file formats including XML, JSON, CSV & other compressed file formats.• Reviewed the HDFS usage and system design for future scalability and fault-tolerance.• Developed Hive Queries in Spark-SQL for analysis and processing the data. Used Scala programming to perform transformations and applying business logic. • Involved in loading data from LINUX file system to HDFS.• Loaded and transformed large sets of structured, semi structured and unstructured data in various formats like text, zip, XML and JSON.• Using GitHub as version control tool for code management. • Extensively using home grown application IBIS to ingest the data from one source to another main component of IBIS is Sqoop.• Defined job flows and developed simple to complex MapReduce jobs as per the requirement.Optimized MapReduce Jobs to use HDFS efficiently by using various compression mechanisms.• As Data ingestion team mainly focused on ingesting data from various databases like Oracle,Mysql,DB2 to Hadoop and Hadoop to AWS.• Extensively using python and pysaprk for the transformations of the data, wrote various glue jobs using pysprak to load the data from Hadoop to AWS S3 in the form a CSV and later loading into AWS Database Redshift.• Wrote some policies on S3 Buckets to load the data after certain period into AWS Snowflake for retrieval in later point of time.• Did a POC on Apache Airflow by writing some DAG’s to schedule the glue jobs we developed in develop environment.• Worked on Mainframe systems to schedule our data loads in production environment.• Used Amazon Web Services (AWS) S3 to store large amount of data in identical/similar repository.