This is Part 2 of a two-part series on AI-powered risk management for program managers. In Part 1, Spotting Risks Is Not Hard. Spotting Them Early Is!, we made the case for change: why traditional risk management struggles to keep pace, and how AI addresses those gaps through four core techniques, namely automated risk identification, predictive modelling, real-time threat monitoring, and scenario planning. This article focuses on implementation. The objective is not to replace your existing risk management framework, but to evolve it into one that is continuous, predictive, and evidence-driven. If you have not read Part 1, we recommend starting there. The link to Part 1 is given below:
Program-Level Risks That Teams Do Not See
Every program draws boundaries: between business lines, between teams, between delivery and governance. Risk data collects cleanly on each side and stops at the line. Each team reads its own side correctly. Nobody is looking at the cumulative effect of the same risk appearing in several places, or at how hard that total hits the program.
Consider a financial services program that carries the same regulatory gap in lending, payments, and operations. Each risk register scores it medium, and each is right. Nobody sums the three, so a concentration risk sits in plain sight, unowned.
An AI layer reads across all three boundaries continuously. That is its structural advantage over any human reviewer. It is not better judgement, but a view that does not stop at the line.
Implementation: A 5-Step Playbook
Figure 1 below shows a pragmatic five-step path, each step building on the last. Step 2 maps directly onto the four techniques from Part 1, turning each from a concept into an operating routine.

Step 1 — Data Foundation. Centralise your risk-related data in a shared, accessible location. For traditional programs, establish a regular export rhythm from your scheduling and issue-tracking systems. For agile and hybrid programs, add retrospective summaries, impediment logs, and dependency registers to your data sources.
Step 2 — Define Approach. For each of the four techniques from Part 1, define a reusable approach and align it to your program’s review rhythm, whether that is a monthly steering committee, a fortnightly sprint review, or a quarterly PI Planning event.
- Automated risk identification: a standard set of inputs you feed the model each cycle, typically your RAID log, recent change requests, and retrospective summaries.
- Predictive modelling: a consistent way of describing your risk variables and asking for scenario outputs.
- Real-time threat monitoring: a defined list of external signals and internal data sources to scan.
- Scenario planning: a standard structure for presenting disruption options and trade-offs to your steering committee.
Step 3 — Tool Stack. Connect your LLM of choice to an automation layer, and establish a routine of feeding external signals, such as regulatory bodies, key vendor news, and industry publications, into it alongside your internal risk data. For agile programs, connect your team-level tools as well. Do this within your organisation’s AI governance policy for data confidentiality and security. If no such policy exists yet, that is the first thing to fix.
Step 4 — Pilot, Baseline, and Train. Scope the pilot before you start. Pick two or three workstreams that share dependencies rather than a single team, since the cross-boundary view is essential, and run for long enough to cover at least two or three of your governance cycles. Baseline your current performance on the metrics in the next section before you begin. Without that baseline you cannot show what AI changed. Then run the AI process in parallel with your existing process. Comparing the two sets of outputs is also a powerful coaching tool, helping program managers, Scrum Masters, and Product Owners see where AI adds value, when to override it, and how to refine inputs over time.
Step 5 — Scale and Govern. Once the pilot demonstrates value against that baseline, scale the process across your program, taking one or two workstreams at a time. Keep a human in the loop for regular review of AI outputs to catch errors, scoring drift, and emerging bias. The realistic barrier is rarely cost; it is data quality and the discipline to keep feeding the process.
Metrics That Matter
Track these against the baseline you captured in Step 4, adapted to your delivery model.
| Metric | How to compute | What good looks like |
| Risk detection lead time | Days between a risk first being identified and its projected impact date. Useful lead time varies by risk class: days for sprint-level impediments, weeks or more for supply chain and regulatory risk. | A rising average, showing the program shifting from reactive to predictive. |
| Risk prediction hit rate | Of all incidents that occurred, the percentage the AI had flagged in advance. | A steady rise, read alongside the false-alarm rate below. |
| False-alarm rate | Flagged risks that never materialised and required no mitigation, as a percentage of all risks flagged. | Falling, or at least stable, while hit rate rises. |
| Mitigation effectiveness | Of the flagged risks you acted on, the percentage mitigated before they caused impact. | Rising in step with hit rate, showing that detection and response are improving together. |
| Variance absorption | Delay in days caused by an unplanned event to the committed delivery date. | Shrinking average delay per event. |
| Steering efficiency | Median days between a risk being escalated and a decision being formally recorded. | A reducing trend, indicating faster and better-informed risk conversations. |
These metrics should be reviewed together. No single metric tells the whole story; the objective is to demonstrate that AI is helping the program detect risks earlier, improve decision quality, and reduce delivery impact.
Challenges and Guardrails
AI is a powerful tool, not a substitute for judgement. Build these guardrails in from the start.
- Hallucinations. Always cross-check high-stakes AI outputs against primary sources before escalating.
- Data quality. Output quality is capped by input quality, so standardise the inputs. Define templates for every report that feeds the process, covering retrospectives, status reports, and risk register entries, and hold teams to them. Inconsistent reporting is a common reason AI risk monitoring produces vague or misleading output.
- Governance. Classify risks by autonomy level: which AI outputs can be acted on directly, which need a human check, and which must go to a steering committee. AI proposes; humans approve anything above the first level. Establish clear accountability, usually with the PMO: when an AI-driven risk decision turns out to be wrong, it must be clear who is answerable for having acted on it.
- Context blindness. AI misreads delivery data without team context. A model flags a sharp throughput drop across two workstreams and drafts an escalation, when the cause is a scheduled site shutdown it could not see. The AI surfaces every candidate signal; the human discards the ones context explains.
- When AI and human judgement disagree. Do not simply overrule and move on. Log the disagreement, the decision, and the outcome, and review the log monthly. Two patterns emerge: inputs that need fixing, and assumptions of your own that the model was right to challenge.
- Fairness and bias. Periodically audit AI outputs for systAgentic Risk Management
- AI agents are emerging that can run continuous, end-to-end risk cycles with limited supervision. They scan waterfall schedules, agile backlogs, and hybrid registers at once, then simulate scenarios, draft mitigation options, and escalate via your communication platform without waiting for the next review meeting.
- This does not relax any of the guardrails above. It raises the stakes on every one of them. A human in the loop catches a hallucination or a misread of team context before it gets used, but an agent acts between reviews, not at them. So the controls have to be working before the agent runs: autonomy classification, clear accountability, and a disagreement log you actually read. Start on low-consequence risk classes.
- Once those are in place, the day starts differently. Before your team sits down, the agent has already reflagged risks that shifted overnight, drafted a briefing note, and updated the RAID log. You begin with decisions rather than information gathering. The note is still a draft, and the decisions are still yours.emic scoring bias, particularly for risks associated with specific teams, vendors, or geographies. Running a second model as a cross-check helps here: models trained differently carry different blind spots, so a risk that only one of them flags is worth a closer look.
- Data privacy. Be mindful of what program data you feed into external LLMs. Commercially sensitive information, personal data, and contractually confidential material should be anonymised or excluded before being shared with any AI tool. Use your organisation’s approved service rather than a personal account.
- Transparency. If you cannot explain why the AI flagged a risk, you cannot confidently act on it. Prefer outputs that cite the specific inputs behind each flag, and require that citation for anything heading to a steering committee.
Agentic Risk Management
AI agents are emerging that can run continuous, end-to-end risk cycles with limited supervision. They scan waterfall schedules, agile backlogs, and hybrid registers at once, then simulate scenarios, draft mitigation options, and escalate via your communication platform without waiting for the next review meeting.
This does not relax any of the guardrails above. It raises the stakes on every one of them. A human in the loop catches a hallucination or a misread of team context before it gets used, but an agent acts between reviews, not at them. So the controls have to be working before the agent runs: autonomy classification, clear accountability, and a disagreement log you actually read. Start on low-consequence risk classes.
Once those are in place, the day starts differently. Before your team sits down, the agent has already reflagged risks that shifted overnight, drafted a briefing note, and updated the RAID log. You begin with decisions rather than information gathering. The note is still a draft, and the decisions are still yours.
Get Started Today
AI-powered risk management is becoming essential for programs facing compressed timelines, regulatory flux, and supply chain fragility, whether they run on waterfall, agile, or hybrid delivery.
You do not need to start at the program level. Pick two or three workstreams in the same program that share dependencies. Take their risk registers, plus last month’s retrospective summaries if any of them runs agile, and give them to your LLM of choice as a single input. Ask what it sees across the workstreams: which risks are related, which look more serious in combination, and what none of them flags on its own. Some of what comes back will already be familiar. What matters is whether anything in it is new, because that is what your current process was not built to surface.
Run it for a few governance cycles alongside your existing process, with your baseline recorded before you start. Expect it to improve as you go: input quality, prompts, and where the AI fits your rhythm all sharpen with use. That is your pilot, and it will show you how to scale to the full program. If your organisation runs several programs, establish the practice fully in one before extending it to the others.
What is the biggest program risk you are managing right now? Share it in the comments and we will suggest an approach tailored to your program
One Response
Is the book out yet sir?