Data Scientist
Current• Developed and optimized a data pipeline using Hadoop and Apache Spark to process and analyze over 2000 healthcare records, which led to the identification of compliance gaps and potential risks, enabling targeted peer assessments and reducing investigation time by 30%, thereby improving regulatory oversight and public safety.• Performed ETL processes in data cleaning, data modeling, data mining and report sharing using Power BI, which reduced turnaround time by at least 70% to meet SLA's standards.• Managed CI/CD pipelines using Azure DevOps to ensure seamless deployment and updates of regulatory compliance monitoring web applications, while creating data-driven dashboards and reports in Tableau to streamline the analysis of over 1000 physician records.• Extracted datasets in Salesforce CRM and loaded the datasets into SQL relational database, created multiple tables required for analysis and leveraged SQL syntax to retrieve, manipulate and extract results of the analysis.• Leveraged Python and Excel Macros to process and analyze unstructured data from over 150 incident reports, which improved the efficiency of extracting actionable insights, reducing manual processing time by 60% and enhancing the ability to identify trends in professional conduct cases.• Implemented a distributed data processing system using Digital Ocean Droplets and AWS EC2 instances to automate the analysis of compliance-related data streams, which led to the processing of 500 real-time events monthly, enabling predictive analytics to identify potential compliance violations, which improved response times by 40% and enhanced regulatory oversight.