Data Mining Intern
โข Developed a real time ETL pipeline using Spark in Scala, which automated and streamlined data processing workflows.โข Extracted financial report data from HDFS (Hadoop Distributed File System) using MySQL; organized and concatenated the data into comprehensive summary tables with Python pandas, and effectively presented the insights to stakeholders.โข Calculated over 600 features using Spark MLlib and evaluated their importance using metrics such as VIS and SHAP.โข Utilized fine-tuned LLM models to automate loan document classification, achieving over 90% accuracy and reducing manual labeling workload by 75%; automated processes using shell scripts and scheduled tasks with cron.