Delivery Confidence Assessment: A Board Guide
In July 2026 the National Infrastructure and Service Transformation Authority published a delivery confidence rating for every major programme the UK government runs. Twenty nine were Green, one hundred and nine Amber, thirty four Red. The numbers made the news. The mechanism behind them deserves more attention, because it is the one governance instrument most private sector boards do not have and could adopt tomorrow.
Ask most boards how their largest transformation programme is performing and they will point to a status report. Ask who wrote it and the answer is almost always the programme itself. That is not governance. It is self-certification with a colour scheme attached, and it is the reason so many boards learn that a programme is in trouble at the point the money has already gone.
What a delivery confidence assessment actually is
A delivery confidence assessment is a rating of whether a programme is likely to achieve what it was funded to achieve, given where it stands today. It is not a measure of activity, spend or milestone completion. It is a forward-looking judgement about the probability of the stated outcome, made against a defined scale, by someone who is not accountable for the result.
The scale government uses is deliberately blunt. Green means successful delivery appears highly likely and no major issues are affecting the programme. Amber means delivery is feasible but significant issues already exist and need management attention. Red means successful delivery appears unachievable against the approved baseline, and that major issues require urgent action if the programme is to proceed successfully.
The value sits in that third definition. A Red rating is not a statement that a programme has failed. It is a statement that it cannot succeed against its current baseline, which is a very different and far more useful thing for a board to hear. It converts a vague sense of unease into a specific question: does the baseline change, or does the programme?
What the 2026 portfolio numbers actually show
NISTA's Major Projects Annual Report 2025 to 2026 covered 189 projects with a combined whole-life cost of £924.2 billion. At the end of March 2026, 15 per cent were rated Green, 58 per cent Amber, 18 per cent Red, with the remaining 9 per cent exempt from a published rating.
The headline most commentators took was that one in five of the government's biggest programmes is rated unachievable. The more instructive reading is the other one. Fifty eight per cent sit in Amber, meaning delivery remains feasible but significant issues already exist. That is the population where intervention is cheap and effective, and it is precisely the population that goes unexamined in most private sector portfolios, because a programme that is broadly on track does not attract board attention until it stops being on track.
Two further details matter more than the ratings themselves. First, forty two projects left the portfolio during the year, twenty six of them having delivered against their objectives, and eighteen projects moved from Amber to Green. Ratings are not a one-way ratchet towards bad news. Second, NISTA has introduced an early warning system that uses existing portfolio data to identify projects at risk of moving to Red before they get there. The direction of travel is from periodic assessment towards continuous signal.
Why private sector boards rarely have an equivalent
Government adopted this discipline because the National Audit Office kept finding the same failure pattern: serious concerns known at working level that never reached the people who could act on them. The pattern is not a public sector characteristic. It is a structural feature of any programme where the person writing the report is measured on the outcome being reported.
In commercial organisations the pressure is often stronger, not weaker. A programme director whose reputation is bound to a go-live date has every incentive to describe a problem as manageable for one more reporting cycle. A systems integrator whose commercial position depends on the current plan has no incentive to tell the board the plan is no longer credible. Neither party is acting dishonestly. They are reporting from inside a position, and no amount of dashboard sophistication corrects for that.
The consequence is that boards receive an assessment of delivery risk from the parties carrying it. In every other domain the board would recognise this immediately. No audit committee would accept management's own opinion of the accounts. Yet programmes representing a material share of the capital budget are routinely governed on exactly that basis.
Four things that make a rating credible
A colour on a slide is not a delivery confidence assessment. Four characteristics separate the two.
- Independence from delivery. The rating must be produced by someone with no accountability for the outcome and no commercial interest in the current plan continuing. This is the single non-negotiable condition, and the one most often compromised when assurance is handed to the integrator or to an internal function reporting through the programme.
- A defined scale, published in advance. The definitions have to be agreed before anyone applies them, so that a rating cannot be negotiated after the fact. If Red can be argued down to Amber in the room, the scale has no meaning.
- Snapshot discipline. A rating describes the position at a point in time and indicates risk at that point. It is not a prediction of ultimate success or failure, and treating it as a verdict is the fastest way to make people fight the rating rather than act on it.
- A forward-looking basis. The assessment must address whether the outcome remains achievable, not whether last month's milestones were met. Programmes almost always report green milestones right up to the point they report a delay, because milestone completion measures the past.
Early warning beats post-mortem
The purpose of rating a programme is not accountability. It is timing. A programme that is going to fail rarely fails suddenly. It accumulates unresolved dependencies, quietly absorbs contingency and rebaselines twice before anyone uses the word recovery. Each of those events is visible months in advance to someone looking for it and invisible to someone reading a status report written by the programme.
NISTA's own framing is that these assessments should not be read as a definitive judgement on whether a project will succeed or fail, but as an early warning that helps identify risks sooner and address issues before they become more serious. That is the right posture for a board too. The rating is not there to allocate blame. It is there to buy time, and time is the only variable in a troubled programme that money cannot replace.
The economics are straightforward. Intervening in an Amber programme usually means adjusting scope, sequence or resourcing. Intervening in a Red programme usually means programme recovery, which costs several multiples more and lands during the period when the organisation can least absorb it. We have set out the cost case for acting earlier in more detail in our view on how independent programme assurance reduces transformation costs.
Introducing delivery confidence assessment to your board
This does not require a governance restructure. It requires four decisions.
Decide which programmes are in scope. Materiality is the right test, not visibility. Any programme whose failure would require a board conversation with shareholders or lenders belongs on the list, however comfortable its current reporting looks.
Agree the scale and write down the definitions before the first assessment. Borrow the government wording if it saves an argument. What matters is that the definitions exist in writing and are not adjusted to suit a result.
Appoint someone independent to produce the rating. This can be a non-executive with the relevant delivery background, an internal function that reports to the board rather than through the programme, or an external reviewer. The test is not seniority. It is whether the person can rate a programme Red without damaging their own position.
Fix the cadence and hold it. Quarterly is sufficient for most portfolios, with the standing expectation that a material change in confidence is reported when it happens rather than held to the next cycle. A rating that only appears when someone requests it is a rating that will not appear when it is most needed.
Common mistakes
- Letting the programme rate itself, then treating the output as assurance.
- Appointing the systems integrator to assess the plan it is being paid to deliver.
- Rating milestone completion rather than the achievability of the outcome.
- Allowing ratings to be negotiated in the meeting, which teaches everyone that the scale is decorative.
- Treating Amber as a comfortable middle. Amber means significant issues already exist, and it is where the cheapest interventions are available.
- Reacting to a first Red rating by challenging the reviewer, which guarantees the second one arrives later.
Where this sits alongside recovery and assurance
Delivery confidence assessment is the diagnostic layer. It tells a board which programmes need attention and how urgently. What follows depends on the answer: reinforced governance and course correction for programmes that remain achievable, and structured intervention for those that are not. The symptoms that typically precede a Red rating are covered in our guide to the signs your IT programme needs recovery, and the portfolio-level application is set out in our note on programme assurance for PE-backed transformations, where the reporting line to the sponsor makes independence particularly valuable.
Intology provides this work independently. We hold no vendor partnerships, take no commissions from software providers or integrators, and have no position to protect in the plan being assessed. That is the whole point. Our Embedded Change Model™ places senior practitioners close enough to the programme to see what is actually happening, while the assessment itself reports to the board rather than through the delivery line.
Government did not adopt published delivery confidence ratings because its programmes are worse than yours. It adopted them because at sufficient scale the cost of finding out late becomes impossible to absorb. Most boards reach that threshold long before they think they have. The instrument is simple, the definitions are already written, and the only genuinely difficult part is accepting a rating you did not want to receive.