- Salary
- $85k – $105k
- Location
- Remote - CO, United States of America
- Workplace
- Remote
- Type
- Full-time
- Department
- Management
- Experience
- 3+ years
- Education
- Bachelor
- Source
- Workday
Description
About Vantage
Vantage powers, cools, protects and connects the technology of the world’s well-known hyperscalers, cloud providers and large enterprises. Developing and operating across North America, EMEA and Asia Pacific, Vantage has evolved data center design in innovative ways to deliver dramatic gains in reliability, efficiency and sustainability in flexible environments that can scale as quickly as the market demands.
Position Overview:
This position will be based remotely in the United States
Site Operations teams manage our buildings under the governance provided by Business Operations, ensuring best-in-class data center services for our customers. Our site teams comprise Critical Facility Engineers who are Vantage's front line for managing critical electrical and mechanical infrastructure operations. Providing our best-in-class services requires in-depth plant equipment knowledge and proficiency in critical facility management processes including Change, Incident, Problem, and Work Management.
Vantage is looking for a resourceful Problem Management Specialist to support the North America Site Operations organization by coordinating the end-to-end Problem Management lifecycle. The role focuses on identifying and eliminating the underlying causes of incidents, through structured root cause analysis, clear ownership, and disciplined follow-through. Serving as a central partner to Site Operations, the Specialist will help ensure that qualifying operational events are converted into complete, factual, and actionable Problem records that strengthen site reliability and reduce the likelihood and impact of recurrence.
Working across Site Operations, the Operations Management Center, Reliability Engineering, Automation Systems, Security, Customer Experience, Legal, vendors, and other technical teams, the Specialist will coordinate root cause analysis tasks, corrective actions, follow on actions and documentation updates while maintaining traceability to the originating incident. The role will guide teams in documenting evidence-based findings, timelines, contributing factors, and support the preparation and governance of accurate internal and customer-facing RCA reports.
A successful candidate will combine strong facilitation, analysis, and operational judgment with the ability to make Problem Management practical for frontline teams. They will maintain high-quality and auditable data, challenge symptom-based conclusions, and translate technical investigations into concise, executive- and customer-ready narratives.
ESSENTIAL JOB FUNCTIONS:
· Manage the end-to-end Problem Management lifecycle for qualifying Site Operations incidents, ensuring Problem records are created, assigned, progressed, and closed in accordance with established standards
· Facilitate structured root cause analysis with Site Operations and technical stakeholders, using incident evidence, event timelines, asset history, procedures, and operational data to distinguish root causes from symptoms and contributing factors
· Coordinate corrective and preventive actions across Site Operations, Reliability Engineering, the Operations Management Center, Automation Systems, Security, vendors, and other support teams, with clear owners, target dates, dependencies, and escalation paths
· Maintain accurate, complete, and auditable Problem, problem task, known error, workaround, outage, work order, and related change information in ServiceNow, with traceability to the originating incident
· Support post-incident follow-up for Site Operations P1 and P2 events by validating operational impact, outage details, investigation evidence, and required Problem Management actions
· Analyze recurring incidents, common failure modes, cause codes, and cross-site trends to identify systemic risks and recommend prioritized opportunities for prevention and continual improvement
· Publish and communicate lessons learned, known errors, workarounds, and validated corrective actions so that improvements are adopted consistently across Site Operations and applicable risks are addressed fleet-wide
· Identify and participate in activities focused on process improvement, automation, tools implementation and QA testing
DUTIES:
· Review qualifying Site Operations incidents and initiate Problem records in ServiceNow when root cause analysis, corrective action, or recurrence prevention is required
· Plan and facilitate root cause analysis sessions, gathering incident timelines, technical evidence, asset history, procedures, monitoring data, and stakeholder input
· Document clear problem statements, business and operational impacts, contributing factors, root causes, cause codes, workarounds, known errors, lessons learned, and recommended actions
· Create, assign, and track problem tasks, corrective and preventive actions, work orders, and related change requests through validation and closure
· Monitor Problem records and assigned actions for aging, service-level performance, overdue commitments, dependencies, and risks; engage owners and escalate barriers as needed
· Prepare accurate internal, executive, and customer-facing root cause analysis reports, status updates, governance metrics, and trend summaries
· Develop and maintain Problem Management procedures, templates, quality standards, dashboards, training materials, and governance routines that make the process practical for frontline teams
· Partner with Site Operations, the Operations Management Center, Reliability Engineering, Automation Systems, Security, Customer Experience, Legal, vendors, and other stakeholders to coordinate investigations and remediation activities
· Define reporting, workflow, automation, and data-quality requirements; support testing and implementation of improvements within ServiceNow and related operational tools
JOB REQUIREMENTS:
· Demonstrated experience applying ITIL Problem Management principles, including problem identification, root cause analysis, known-error management, corrective-action tracking, and formal closure
· Bachelor of Science degree in Information Technology, Business Management, related field, or equivalent experience
· 3 years of experience in Problem Management or equivalent supporting function
· Hands-on experience managing Problem records, problem tasks, known errors, workarounds, and related incident and change records in ServiceNow or a comparable IT service management platform or computerized maintenance mananagement system (CMMS).
· Ability to facilitate structured investigations with technical and operational teams, challenge symptom-based conclusions, and guide stakeholders toward evidence-based root causes and practical corrective actions
· Working knowledge of root cause analysis methods such as the 5 Whys, causal-factor analysis, fault-tree analysis, fishbone diagrams, or equivalent structured techniques
· Strong organizational, written and verbal communication skills, with experience coordinating multiple action owners, dependencies, due dates, service-level commitments, and escalations across concurrent investigations
· Ability to translate complex technical findings into concise, factual, and audience-appropriate root cause analysis reports, executive updates, customer communications, and lessons learned
· Proven ability to work effectively with cross functional teams, engineering, service management, vendors, and business stakeholders while maintaining objectivity, accountability, and a constructive approach during complex or high-visibility investigations
· Advanced skills in Microsoft Office 360Suite – Excel, Word, Power Point, Project, and Visio
· Data Center, high-tech, or rapid growth industry experience is strongly preferred, but not required
Up to 5% travel required may change if business needs expand.
Additional Details
- Salary Range: $85,000-$105,000 Base + Bonus (this range is based on Colorado market data and may vary in other locations)
- This position is eligible for company benefits including but not limited to medical, dental, and vision coverage, life and AD&D, short and long-term disability coverage, paid time off, employee assistance, participation in a 401k program that includes company match, and many other additional voluntary benefits.
- Compensation for the role will depend on a number of factors, including your qualifications, skills, competencies, and experience and may fall outside of the range shown.
#LI-WW1
#LI-Remote
We operate with No Ego and No Arrogance. We work to build each other up and support one another, appreciating each other’s strengths and respecting each other’s weaknesses. We find joy in our work and each other, actively seeking opportunities to inject fun into what we do. Our hard and efficient work is rewarded with an above market total compensation package. We offer a comprehensive suite of health and welfare, retirement, and paid leave benefits exceeding local expectations.
Throughout the year, the advantage of being part of the Vantage team is evident with an array of benefits, recognition, training and development, and the knowledge that your contribution adds value to the company and our community.
Don't meet all the requirements? Please still apply if you think you are the right person for the position. We are always keen to speak to people who connect with our mission and values.
Vantage is an Equal Opportunity Employer.
Vantage does not accept unsolicited resumes from search firm agencies. Fees will not be paid in the event a candidate submitted by a recruiter without an agreement in place is hired; such resumes will be deemed the sole property of Vantage.
We’ll be accepting applications for at least one week from the date this role is posted. If you're interested, we encourage you to apply soon—we’re excited to find the right person and will keep the role open until we do!