Site Reliability Engineer
Current• Worked as an SRE for AppDynamics Cisco managing Infrastructure and maintaining the SLA. • Assist in development of custom dashboards and Business Transaction auditing for efficient monitoring of key business applications• Support and maintain Build and Release Engineering Public/Private cloud Infrastructure.• Responsible for management and administration of AWS, TeamCity, and other automation infrastructure.• Serve as an escalation point for production issues during shift or as required.• Managing AWS EKS clusters versions up to date by upgrading the cluster and node group using Terraform and helm.• Integrated vulnerability scanning tools such as SonarQube and Black Duck into TeamCity CI/CD pipelines to ensure code security and compliance.• Conducted regular audits of code quality and dependency vulnerabilities, providing recommendations to development teams for proactive fixes.• Ensured high availability and minimal downtime by architecting disaster recovery solutions for mission-critical applications using Kubernetes, Helm and TeamCity.• Managing Multiple TeamCity Ephemeral and Hardened Build Agents which supports AIX, Linux, MAC, Windows and HP-UX.• Provisioning testkube instance setup and debugging sessions for product teams to validate new features of the application.• Managing Kotlin files for TeamCity to manage a Library based pipelines which will be used by platform teams.• Work with key Business stakeholders to understand their business requirements, recommend potential solutions, and secure resources to deliver• Maintain key operational metrics and provide regular updates to upper management• Deliver operational services that focus on empowering our employees and reducing cost per end-user.• Seek opportunities to streamline standard operating procedures through automation.