Software Engineering Manager
CurrentOverseeing the implementation and maintenance of monitoring systemsto ensure the health and performance of IT infrastructure, applications,and servicesDefining alerting thresholds, configuring alerting systems, and managingalert notifications to ensure timely response to incidents.Leading incident response efforts by coordinating with cross-functionalteams to diagnose and resolve issues impacting system availability andperformance.Developing and implementing strategies to enhance observability,including logging, metrics, and tracing, to gain insights into systembehavior and improve troubleshooting capabilities.Assessing and selecting monitoring and observability tools, overseeingtheir deployment, and ensuring they meet the organization's requirements.Analyzing system performance metrics and logs to identify trends,anomalies, and areas for optimization.Documenting monitoring processes, procedures, and best practices, andproviding training to team members on observability tools and practices.Continuously evaluating and improving observability practices, tools, andprocesses to enhance system reliability, availability, and performance