Senior Software Engineer
CurrentBig Data and Cloud Computing* Automation of Cluster (HDP) construction, 2 topologies are proposed:- Integration: One Master node, N Data nodes, Analysis nodes, Edge nodes- Production: Five Master nodes, N Data nodes, Analysis nodes, Edge nodes- Each cluster has his own LDAP branch and his own KERBEROS realm- Business user provisioning- Installation of analytics tools on Analysis nodes (hue, jupyter, rstudio)- Installation of ETL tools on edge nodes (datastage, talend, dataiku)* Automation of Cluster scaling, patching, and health checking:- Cluster Scaling (add/remove) of data nodes, analysis servers, and edge nodes- Security patching and upgrade of Ambari & HDP stack- Cluster health check (all hosts, services & components) and also connectivity of tools in analysis & edge nodes* Development of backup strategies for HDP Clusters:- Stretched topology over two DC- Cross-Realm between primary and secondary clusters and scheduling of distcp jobs- Data replication using hdfscli (delta based on Snapshot diff in primary cluster)* Big Data IT rules & best practice for better use of the cluster resources* Env: HDP 2.6, HDP 3.1, SPARK, HIVE, HBASE, YARN, HDFS, ANSIBLE, JENKINS, GIT/GITLAB, NEXUS, PYTHON, SHELL, DATASTAGE, TALEND, ORACLE EXADATA, HUE, JUPYTERHUB, RSTUDIO, ELASTICSEARCH, LOGSTASH, KIBANA, MONGODB, CASSANDRA, ORIENTDB, NEO4J