Site Reliability Engineer
CurrentResponsible for Operating, maintaining and administering solutions that contribute to the operational efficiency, availability and visibility of customer infrastructure.Planning maintenance activity, design documentation and standard procedures.Provide Root Cause Analysis reports for outages/incidents (ITIL - Problem Management).Observe and provide feedback on the current state of the client’s infrastructure, and identify opportunities to improve resiliency, reduce the occurrence of incidents and automate repetitive administrative and operational tasks.Responsible for improving and maintaining team documentation about client systems and infrastructure, procedures, policies and schedules.- Windows Server deployment, configuration and performance tuning.- Active Directory architecture, deployment, migration and group policy management.- TCP/IP networking, NIC teaming, and network services configuration (DNS, NTP, DHCP, etc.).- Strong understanding of backup solutions and how to map requirements into solutions.- Strong understanding of multiple monitoring solutions and how to implement and manage them.- Microsoft clustering technologies.- Scripting and automation of administrative tasks using PowerShell.- Cloud automation and Scripting (Azure Automation, PowerShell, Python, YAML)- Cloud Networking.Knowledge of VPC, Virtual networks, Resource groups- Strong understanding of routing and dynamic routing protocols (BGP) in a cloud capacity and/or using Cisco/Juniper equipment.- Linux operating systems administration, bash and/or shell scripting).- Experience with planning and executing server migrations, P2V conversions, on-prem to cloud migration strategy.