Data Scientist
CurrentApplying advanced data science techniques to network analytics in support of effort to reduce false positives in cybersecurity application. Responsible for data ingest (combination of web-scraping, APIs, text files and CSV). Processing has included Latent Dirichelet Allocation and K-means clustering (unsupervised machine learning methods), Natural Language Process (NLP) using Natural Language Toolkit (NLTK), and logistic regression (machine learning method for categorization). Have implemented graph based structures of data in neo4j and more recently, networkx. Processing pipelines supported by Python (including a full set of data science tools - most especially pandas/numpy and scikit-learn), R, general Linux tools such as Perl and Emacs. Much of my time is spent on 'feature engineering' to allow for meaningful application of machine learning tools. Core data elements are persisted in MySQL.Have contributed to enhancements to python module ipwhois. Seeking other opportunities to support interesting open source projects.Migrating some of my workflows from scripts to Jupyter Notebooks - great for self-documenting, reproducible analyses.Enjoy Python a lot, attended Pycon2018 in Cleveland. I write __iter__ often, __enter__ and __exit__ not as much :)