Data Scientist
Current• Lead the full machine learning system implementation process: collecting data, model design, feature selection, system implementation, and evaluation.• Utilized Spark, Scala, Hadoop, HBase, Spark Streaming, MLLib and R a broad variety of machine learning methods including classifications, regressions, dimensionally reduction etc.• Creating various B2B Predictive and descriptive analytics using R and Tableau.• Installing NVIDIA Drivers along with CUDA 9• Used text mining and NLP techniques find the sentiment about the organization.• Developed unsupervised machine learning models in the Hadoop/Hive environment on AWS EC2 instance.• Worked with datasets of varying degrees of size and complexity including both structured and unstructured data.• Responsible for Installing Tensorflow for GPU• Participated in all phases of data mining, data cleaning, data collection, developing models, validation, visualization and performed gap analysis.• Data wrangling to clean, transform and reshape the data utilizing Numpy and Pandas library.• Data Storyteller, Mining Data from different Data Source such as SQL Server, Oracle, Cube Database, Web Analytics, Business Object and Hadoop. Provided AD hoc analysis and reports to executive level management team.• Contributed to data mining architectures, modeling standards, reporting, and data analysis methodologies.• Worked with different sources such as Oracle, Teradata, SQLServer and Excel, Flat,• Complex Flat File, Cassandra, MongoDB and HBase files.• Conduct research and made recommendations on data mining products, services, protocols, and standards in support of procurement and development efforts.• Used Python, R and SQL to create Statistical algorithms involving Linear Regression, Logistic Regression, Random forest, Decision trees for estimating the risks.• Developed statistical models to forecast inventory and procurement cycles.