Source description
About the role
Group Company: MINDSPRINT DIGITAL (INDIA) PRIVATE LIMITED
Designation: Lead NoC
Office Location: Chennai
About the Role
You’ll lead NOC operations ensuring high availability of enterprise infrastructure—network, servers, storage, virtualization, cloud, and security platforms. You will oversee incident response, major incident management (MIM), proactive monitoring, change governance, and continuous improvement of monitoring/alerting processes with a strong focus on SLA compliance and service reliability . Key Responsibilities Operations & Monitoring Own 24×7 monitoring of infrastructure components: WAN/LAN, firewalls, load balancers, VPN, proxies, servers (Windows/Linux), virtualization (VMware/Hyper-V), storage (SAN/NAS), cloud (AWS/Azure/GCP), and backup systems. Maintain and tune monitoring tools (e.g., SolarWinds, Zabbix, PRTG, Nagios, Dynatrace, AppDynamics, SCOM, Splunk/ELK) for accurate alerting and noise reduction. Implement and optimize observability (metrics, logs, traces) and health dashboards. Incident & Problem Management Lead L2/L3 incident response , triage, RCA, and restoration; coordinate across Network, Server, DB, Security, and Application teams. Act as Major Incident Manager during SEV1/SEV2 events, managing bridge calls, stakeholder communication, and war-room resolution. Drive Problem Management : document RCAs, implement corrective/preventive actions (CAPA), and track trend analytics to reduce recurrence. Change, Release & Compliance Review/approve change requests (CAB participation), ensure robust pre/post validations, rollback plans, and impact assessments. Ensure compliance with ITIL processes , internal controls, audit requirements, and security baselines (patching, backups, DR drills). SRE & Automation (Nice-to-Have / Preferred) Enhance reliability through runbooks , self-healing scripts , and automation (Python/PowerShell/Bash/Ansible). Contribute to error budgets, SLIs/SLOs, and post-incident reviews . Stakeholder Management & Reporting Provide real-time status updates to business/leadership during incidents; maintain a consistent communication cadence. Publish daily/weekly/monthly operations reports : availability, MTTR/MTTA, incident trends, change success rate, capacity/utilization, and SLA adherence. Mentor and lead a small NOC team (5–15 engineers), oversee training, shift planning, and performance. Required Skills & Experience Technical
Networking: TCP/IP, BGP/OSPF, MPLS, SD-WAN, switching/routing, VPN, DNS/DHCP, QoS. Security: Firewalls (Palo Alto/Checkpoint/FortiGate), IDS/IPS, WAF, proxies, SSL/TLS, certificate management. Systems: Windows/Linux administration, AD/Group Policy, patching, backup/restore, job schedulers. Virtualization & Storage: VMware/ESXi/vCenter, Hyper-V, SAN/NAS concepts, snapshots, replication. Cloud: AWS/Azure/GCP fundamentals (VPC/VNet, security groups/NSGs, load balancing, monitoring). Monitoring/Observability: SolarWinds, Zabbix, PRTG, Nagios, SCOM, Dynatrace/AppDynamics, Splunk/ELK, Grafana. Scripting/Automation: PowerShell, Bash, Python; knowledge of Ansible/Terraform (bonus). ITSM Tools: ServiceNow/Jira/Remedy for Incident/Change/Problem/CMDB
Required abilities
Physical:
Other:
Work Environment Details:
Specific requirements
Travel:
Vehicle:
Work Permit:
Other details
Pay Rate:
Contract Types:
Time Constraints:
Compliance Related:
Union Affiliation:
More at MINDSPRINT INC