Problem Manager – ITIL, ServiceNow, Root Cause Analysis
Job Description:
- Lead enterprise-wide Problem Management, Root Cause Analysis, and Incident Prevention initiatives
- Facilitate 5 Whys RCA sessions for P1 incidents and recurring P2 incidents
- Review logs, monitoring tools, dashboards, and change records to establish timelines and contributing factors
- Use AI-assisted tools for investigations, timeline generation, RCA documentation, and knowledge capture
- Document root causes and corrective actions in ServiceNow Problem Management records
- Lead the Problem Management lifecycle from investigation through closure
- Partner with technical teams to create Remediation Action Plans
- Track remediation progress, escalate risks, and enforce stakeholder accountability
- Validate solutions and measure service reliability, operational risk, and incident recurrence improvements
- Identify recurring patterns across infrastructure, cloud, application, and network incidents
- Recommend monitoring, automation, and architectural improvements
- Facilitate weekly Problem Review meetings
- Maintain executive dashboards and Problem Management reporting
- Deliver executive-ready summaries, status updates, and RCA reports for customers and senior leadership
- Improve governance processes, operational standards, runbooks, and reporting practices
- Champion automation and AI-enabled workflows
Requirements:
- 7+ years of experience in IT Operations, Problem Management, Incident Management, Major Incident Management, or related disciplines
- Strong understanding of ITIL Problem Management, Incident Management, and operational governance
- Experience conducting Root Cause Analysis, 5 Whys investigations, and post-incident reviews
- Hands-on experience with ServiceNow or comparable ITSM platforms
- Strong analytical, documentation, facilitation, and stakeholder management skills
- Ability to influence cross-functional teams and drive accountability without direct authority
- Experience creating executive-level communications, dashboards, and operational reporting
- Broad technical understanding of Windows and Linux platforms
- Knowledge of VMware and Hyper-V virtualization
- Performance troubleshooting across CPU, memory, storage, and I/O
- Knowledge of SAN and NAS environments, storage performance, redundancy, and resiliency
- Knowledge of routing, switching, VLANs, DNS, firewalls, load balancing, and packet-flow analysis
- Ability to interpret network latency and packet loss
- Knowledge of AWS and Microsoft Azure, including cloud networking, compute, identity, and storage services
- Understanding of cloud architecture and failure-mode analysis
- Understanding of APIs, microservices, distributed systems, CI/CD, and deployment pipelines
- Understanding of application defects, configuration drift, and dependencies as causes of operational incidents
- ITIL Foundation or higher certification preferred
- Experience supporting large-scale enterprise environments preferred
- Experience with operational analytics, trend analysis, and KPI reporting preferred
- Exposure to automation, AI-enabled workflows, or AIOps practices preferred
- Experience with executive stakeholders and client-facing incident communications preferred
Benefits:
- Bonus or incentive eligibility based on business need
- Health insurance coverage
- Voluntary dental and vision programs
- Life and disability insurance
- Retirement savings plan
- Paid holidays
- Paid time off (PTO) or vacation and/or sick time