Service Manager, Global · Workday
XavierBennet
14+ years leading major incident, crisis and 24x7 command-centre operations across telecommunications, SaaS and critical infrastructure.

When systems fail, people look for answers. My role is to find them, restore service, and make the next failure less likely.
Who am I
Summary
Senior IT Operations Leader with 14+ years owning major incident management, 24x7 Command Centre / NOC operations, and critical service reliability. Expert at navigating complex, multi-supplier ecosystems to ensure rapid service restoration; ran shift teams of ~40 engineers at Optus and currently handle high-stakes global P1/P2 incidents at Workday. Proven track record of protecting critical customer journeys and advising C-level executives through national-news-level events, while leveraging AIOps (Gemini, Claude) to optimize real-time stakeholder communication and automate trend analysis.
Key achievements
Three nights — and years — that show how I run an incident, a backlog, and a clock.

40 minNational restore
01 · Optus
Sixty percent of a national network
NOC shift lead for ~40 engineers. Full service restored in 40 minutes.
How I ran it
8 → 2Comms, minutes
02 · Workday
Eight minutes to two
Agentic incident comms on Gemini, PagerDuty and Slack — plus a Claude runbook agent.
How I ran it
−56%Problem tickets
03 · Tabcorp
Ninety-seven problems down to forty-two
56% fewer problem tickets in eleven months, under multi-state wagering regulation.
How I ran it
Impressive work bringing the network back on in such a short duration.
What I bring
Major incident command, crisis, problem discipline, and AIOps — earned in telco, data centres, wagering and global SaaS.
Major incident command
End-to-end ownership of P1/P2 restoration: engineering, vendors, executives, and the clock. I have run 24x7 NOC shifts of ~40 engineers and global SaaS bridges in the same week of a career.
Crisis and disaster response
Bushfires, national network events, Christmas traffic, iPhone launches. Onshore, offshore and field teams, plus agencies — TIO, RFS, emergency services — and ministerial briefings when the impact is public.
AIOps that actually ships
Not a Copilot tab. Agentic workflows on Gemini, Claude, PagerDuty and Slack that cut broadcast time from eight minutes to two, parse MTTD/MTTI/MTTR, and walk a team through the runbook for the thing that is actually down.
Problem and PIR discipline
Root cause, known-error database, trend analysis, and post-incident reviews that executives will sit through. Recurring noise gets named, owned, and reduced — not parked in a ticket pile.
Multi-vendor governance
SLAs, service reviews, action trackers, and the nerve to renegotiate when post-incident data shows the vendor is the weak link. Telecom, data centre, cloud and wagering suppliers included.
C-level and regulatory rooms
Clear, timed communication under pressure — Fortune 500 SaaS, ASX operators, and state wagering regulators. Confidentiality when the incident is a security vulnerability. Evidence when the government asks.
Tools I actually run
PagerDuty
Jira
ServiceNow
Slack
Google Gemini
Claude
Cursor
Where I have sat
Optus, Nokia, Macquarie Technology, Tabcorp, Workday. Same job in different clothes: restore the service, tell the truth, make the next one quieter.
Workday

May 2024 — Present
NASDAQ: WDAY. Enterprise AI platform for people and finance — more than 11,500 organisations, including over 60% of the Fortune 500. About 20,000 employees.
Service Manager, Global · North Sydney
Global P1/P2 governance, APAC crisis lead, and AIOps on a platform used by more than half of the Fortune 500.
- Owned end-to-end major-incident restoration across engineering, vendors and the business, including executive-only bridges for a high-priority security vulnerability.
- Owned problem management end to end — intake, trend analysis, RCA quality, known-error records, and follow-through with engineering and vendors until the action actually closed.
- Ran service management end to end: service reviews, action tracking, and owners with dates so outstanding work was visible and closed, not parked in a slide.
Tabcorp

Feb 2023 — Apr 2024
ASX: TAH. Australia’s largest wagering and gaming operator — TAB retail and digital, racing and sports. About 3,000 people, listed in Melbourne.
Major Incident and Problem Manager · Sydney
Incident, problem and change for Australia’s largest wagering operator, under multi-state regulatory oversight.
- Reduced problem tickets 56% (97 → 42) in eleven months through trend analysis and vendor risk acceptance.
- CAB lead for the weekly operations change call; cut change lead times 27% and ended restore-order disputes mid-outage.
- Closed the regulatory reporting backlog by opening a working channel with each state authority.
Macquarie Technology
Jun 2021 — Jan 2023
ASX: MAQ. Australian data centre, cloud, cyber and telecom for business and government — not the bank. Headquartered in Sydney, roughly 500–1,000 people.
Major Incident and Problem Manager · Sydney
Major incident ownership across cloud and data-centre estates, including events that hit up to 200 SME customers.
- Introduced a fatigue rule on the 24x7 command centre: engineer swap after six hours on a critical, so the bridge gets a fresh mind instead of a spent one.
- Used post-incident metrics to expose a vendor service flaw, renegotiated the contract, and drove 12% cost savings.
- Ran monthly customer and vendor SLA reviews with an action tracker that made outstanding issues visible.
Nokia

Jun 2018 — May 2021
Finnish network vendor. About 78,000 people worldwide. Builds the radio and core kit behind operators including Optus and Vodafone in Australia.
Service Delivery Manager — Incident, Problem & Change · Sydney
Onshore owner for offshore incident, problem and change teams in India and the Philippines, plus crisis work for two mobile network contracts.
- Led joint Optus–Vodafone resource sharing on JV sites during natural disasters — faster restore, lower cost.
- Trained the offshore major-incident team; the group took a quarterly performance award. Built macros so tier-1 could recognise a major without waiting for onshore.
- Designed and trained a Python disaster-management tool that improved site-list accuracy by up to 17%. Bushfires, Christmas traffic, iPhone launch — RFS and telco authorities on the same plan.
Optus

Dec 2012 — May 2018
Singtel’s Australian mobile and fixed network. The country’s second-largest telco, about 8,000 people, national coverage.
Incident Controller / Duty Manager · Macquarie Park, Sydney
24x7 NOC shift lead for ~40 engineers. Central authority on major outages, then the promotion into that seat.
- Coordinated recovery of a major that impacted ~60% of the national network; full service in 40 minutes.
- Led bushfire response on critical network infrastructure with onshore, offshore and field teams, TIO and emergency services, plus executive and ministerial briefings.
- Promoted from incident commander to NOC shift lead in two years. Recognised for back-to-back high-severity bridges lasting 14+ hours. Stood up a structured CAB that earned C-level recognition.
How I write in the room
Non-confidential samples — the skeleton I use for a P1 broadcast, a crisis brief, and a PIR. No customer data. Click to open.
P1 comms template
What happened, who is hit, next action, next clock. The sentence executives can take upstairs.
Open templateCrisis / DR brief
One page for the night the event is nationally significant. Situation, actions, decisions, next update.
Open templatePost-incident review
Empty PIR. Impact, timeline, 5 Whys, clocks, comms review, actions with owners.
Open template
How I work
- 01
Restore first. Then tell the truth.
The job on a P1 is service back, then an honest timeline. Executives do not need theatre. They need impact, next action, and when the next update lands.
- 02
The same failure twice is a process problem
Incident command without problem management is just heroics. Trend the majors, train the room, fix the runbook. That is how you cut volume instead of getting better at firefighting.
- 03
Tools are not competence
PagerDuty, Jira, Gemini and Claude are useful when they change MTTD and MTTR. A licence on a slide is not operations. I build the workflow, then I measure it.
- 04
People last longer than the bridge
Fourteen-hour bridges happen. Fatigue rules, engineer swaps after six hours, and a culture that can say the shift is spent — that is how you still have a team tomorrow.
Before the interview
The questions I get asked in the first ten minutes.
- Where do you work from?
- Sydney. Not a relocation story.
- Will you actually take a 2am bridge?
- Yes. I ran 24x7 NOC shifts of ~40 engineers. Fatigue rules exist so the room still has a team tomorrow. I am not interested in a job that pretends incidents happen at 10am.
- What are you looking for?
- Major incident, problem, crisis, command-centre leadership. Global or national P1/P2 ownership. If the role is a ticket queue with no authority on the bridge, it is not the job.
- Permanent or contract?
- Permanent first. Contract if the seat is real — IC / problem / crisis with executive air cover — not a backfill with no mandate.
- Can we see real RCAs and PagerDuty?
- No. Those belong to the employer. The Playbooks section is the skeleton I use, stripped of customer data. In the room I can walk a PIR from memory without opening a ticket.
- Are you only telco?
- No. Telco, data centre, wagering, global SaaS. Same job in different clothes: restore the service, tell the truth, make the next one quieter.
Start a conversation
Send a short note. It lands in my Gmail — Bennet.rk@gmail.com — and I reply from there.
- Bennet.rk@gmail.com
- linkedin.com/in/xavierbennet
- Based
- Sydney, Australia








