Data Analyst
• Transformed Hive/SQL queries into Spark transformations, using Spark RDDs, Python, and Scala, improving data processing speed by 30% and enhancing the analytics capabilities of the team.• Implemented Slowly Changing Dimension (SCD) transformations in ETL processes, maintaining historical data accuracy in the data warehouse, which resulted in a 20% improvement in data integrity over time.• Configured a Lambda function to trigger from S3 bucket events, streamlining data ingestion pipelines and reducing data processing latency by 25%.• Led ETL testing activities, ensuring data integrity from extraction through to loading into the data warehouse, which improved data quality by 40%.• Deployed code through Jenkins and managed version control with Git, enhancing team collaboration and reducing deployment errors by 15%.• Analyzed product effectiveness using customer feedback, performing data collection, cleaning, and pre-processing. Utilized SQL and Excel for quantitative analysis, uncovering insights that informed a 10% increase in product enhancements.• Collected and implemented business requirements from stakeholders, achieving 100% deliverable compliance with project specifications and enhancing stakeholder satisfaction.• Imported data into Cassandra using Sqoop, optimizing data availability and reducing data transfer times by 20%.• Extracted data from SQL Server to Flat Files/Excel, streamlining reporting processes and improving data accessibility for analysis by 30%.• Managed data migration processes, ensuring data integrity and quality, which contributed to a 25% reduction in data-related issues post-migration.