Mts
Current- Design and development of High Availability subsystem that provides resiliency to data node failures in distributed cluster. HA subsystem manages cluster quorum service(zookeeper), storage pool services within a data node, NVRAM replication and deduplication index replication across controllers in different data nodes. Leverages heartbeats through disks, cluster quorum service, NVRAM replication service and deduplication index replication service to maintain redundancy and provide fault tolerance in the cluster without compromising data integrity. - Distributed master and agent framework that coordinates the distributed software upgrade of data node controllers and compute nodes by closely interacting with High Availability subsystem.- Seamless addition of new node in the cluster using zero configuration networking. Used the same zero config technology to support network config using single pane of glass.- Involved in design and development of Adaptive Pathing to provide improved aggregated network bandwidth and increased end-to-end network availability.- Health monitoring, events/alarm generation for HA subsystem and other components in hardware platform.- New platform bringup (kernel bring up, firmware upgrade, new model handling).- Mentored junior engineers and interns in High Availability and platform teams.- Developed various tools to aid debugging and root causing issues in High Availability and networking areas.