Modernisation business cases have become more sophisticated in recent years. Most now account for architectural debt, the accumulated cost of tightly coupled systems and outdated platforms, and many include a view of data debt, recognising that fragmented and poorly governed data will slow any programme that depends on it. Far fewer consider a question that often proves decisive once work begins: can the organisation reliably confirm what its existing systems do today, and prove that they still do it after change?
That question sits at the heart of what is described as test debt. Legacy systems frequently depend on manual regression cycles, undocumented business rules and a small number of long-serving specialists who know where the edge cases lie. While the system runs unchanged, these weaknesses stay largely invisible. Once modernisation begins, they surface quickly, in the form of delayed discovery, rework during integration and heightened risk at release. In our experience, test debt is one of the most common reasons that modernisation programmes take longer and cost more than their business cases predicted.
The wider cost of weak software quality is considerable. The Consortium for Information and Software Quality estimated that poor software quality cost the US economy at least $2.41 trillion in 2022, with accumulated technical debt reaching around $1.52 trillion. Test debt is a significant contributor to both figures, and it is measurable. Leaders who ask for the right evidence before approving a modernisation route can use it to choose between options, set realistic timelines and fund assurance work at the point where it delivers most value.
What test debt is
Test debt is the gap between the assurance an organisation needs to change a system safely and the assurance it can actually provide. It accumulates gradually, often over many years, as systems are extended faster than their tests, as testers move on, and as manual processes become the accepted way of confirming that a release is safe.
It tends to take seven forms:
1. Unknown behaviour: where business rules exist only in code or in the knowledge of a few individuals, and nobody can state with confidence what the system does in every circumstance. The UK government's own guidance on legacy technology recognises this risk, listing too few people with the required knowledge and skills among its indicators of a legacy system. McKinsey has made a similar observation, noting that the programmers who built and maintain many ageing enterprise systems are now reaching retirement age.
2. Workaround-driven process: where bugs or idiosyncrasies in a legacy system begin to dictate the business process, with people developing workarounds that gradually become the accepted way of working, until nobody can say with certainty what the intended process should be.
3. Manual-only regression: where every release depends on people working through test scripts by hand, which limits how often change can happen and how much of the system can be checked each time.
4. Brittle or untrusted automation: where automated tests exist but fail intermittently or cover the wrong things, so teams have learned to disregard their results.
5. Missing test data: where realistic data for testing is unavailable, out of date or unusable for reasons of privacy.
6. Environments that do not reflect production: so that tests pass in conditions that differ materially from the ones the system will face.
7. Absent non-functional baselines: where nobody has recorded how the current system performs under load or how it behaves under attack, leaving the modernised system with no benchmark to meet.
Most organisations carry several of these forms at once, and they compound one another. A system with unknown behaviour and no realistic test data, for example, cannot be characterised quickly, however capable the team assigned to it.
How test debt creates hidden cost
Test debt rarely appears as a line in a modernisation budget, yet it shapes cost and risk at every stage of delivery.
It first appears at discovery. Before a team can modernise or replace a system, it must understand the behaviour that needs to be preserved. Where that behaviour is undocumented and untested, discovery takes longer, relies more heavily on scarce individuals and produces a less reliable picture of scope. Estimates built on that picture carry more uncertainty than they appear to.
It then appears at integration. As new components replace old ones, teams need confidence that the combined system still behaves as the business expects. Without trusted regression coverage, each integration becomes a period of investigation, and defects that could have been caught early surface late, when they are most expensive to resolve.
It appears most visibly at release. A modernised system that cannot be tested at realistic volume, in a representative environment, with meaningful data, will reach production carrying risks that nobody has been able to measure. The consequences at that point extend well beyond the technology team to customers, regulators and the organisation's reputation.
The effect varies by modernisation route. Our companion article on the rebuild decision sets out six routes for any capability, and test debt affects each differently. A rebuild depends heavily on a complete understanding of existing behaviour, which makes it the route most exposed to test debt. Incremental modernisation is more forgiving, since each stage can be tested against the part of the system it replaces, provided that characterisation of that part is possible. Even a decision to buy a packaged product depends on knowing which existing behaviours the new product must reproduce. Test debt is therefore a direct input to the risk and capacity for change assessments that any route decision should include.
The arrival of AI-assisted development adds a further dimension. AI tools can accelerate the writing and refactoring of code considerably, but Google's DORA research found that higher AI adoption is associated with increases in both delivery throughput and delivery instability. Faster change places more weight on the ability to verify that change, so organisations planning AI-assisted modernisation have even stronger reason to understand their test debt before they begin.
Lessons from TSB
The 2018 migration at TSB remains one of the most thoroughly documented examples of what can happen when assurance falls short on a major programme. The bank moved its customers onto a new core banking platform in a single event, and while the data migrated successfully, the platform experienced technical failures immediately. All of TSB's branches and a significant proportion of its 5.2 million customers were affected, and some issues persisted until December 2018.
The independent review commissioned by the bank's board and carried out by Slaughter and May identified several areas that could have been handled differently, including the need for stronger supplier oversight and questions about how testing was carried out. The regulators later fined TSB £48.65 million for operational risk management and governance failures.
TSB's circumstances were unusual in their scale and complexity, and the causes of the disruption extended well beyond testing. However this illustrates a principle that applies to modernisation programmes of every size in that the confidence of a go-live decision can only be as strong as the assurance evidence behind it.
The assurance evidence required
Leaders approving a modernisation route are rarely in a position to assess test coverage line by line. They can, however, ask for evidence that reveals the scale of test debt and its likely effect on the programme. We recommend seven areas of evidence, each paired with a question a sponsor can reasonably ask at a steering meeting.
| Evidence | Question for the steering group |
|---|---|
| Map of critical business behaviours and their test coverage | Which behaviours matter most, and how many can we currently prove? |
| Characterisation tests of current behaviour | Have we captured what the system does today, including its exceptions? |
| Regression cycle time | How long does it take to confirm that a release is safe? |
| Defect escape rate | How many defects currently reach production, and how serious are they? |
| Test data availability | Can we test with realistic, lawful data at production volume? |
| Environment parity | How closely do our test environments match production? |
| Performance and security baselines | What standard must the modernised system meet, and how will we know it has? |
Taken together, these seven areas give a sponsor a clear picture of how much of the system's behaviour is known, how quickly change can be verified and how closely testing conditions match reality. Where the answers are weak, the programme's timeline and budget should reflect the work needed to strengthen them, and the choice of route may need to change.
An illustrative scenario shows how this evidence can alter a plan. Consider an organisation preparing to replace a claims processing system. The initial business case assumes a twelve-month rebuild. An assurance review before approval finds that only a third of the system's critical behaviours are covered by any form of test, that regression depends on a three-week manual cycle and that no anonymised test data exists at realistic volume. With that evidence, the sponsor can make a more informed choice: fund a short period of characterisation and test data work before build begins, and move from a single rebuild to a staged modernisation in which each component is replaced and verified in turn. The overall timeline may lengthen modestly, while the risk carried into release falls substantially.
Paying down test debt before build begins
Test debt can be reduced deliberately and efficiently, and the most effective point to do so is before significant build work starts.
Characterisation testing is usually the first step. The technique, described by Michael Feathers in Working Effectively with Legacy Code, involves writing tests that record what a system currently does, whether or not that behaviour matches the original specification. These tests create a safety net against which change can be measured, and they often surface undocumented business rules that would otherwise emerge late in delivery.
Targeted automation follows. Automating every test is rarely necessary or economical. Concentrating automation on the behaviours that carry the greatest business value and risk provides the greatest return, and it shortens regression cycles where speed matters most.
A test data strategy completes the foundation, covering how realistic data will be generated, anonymised and refreshed across environments. This is often the longest-lead item, particularly in regulated sectors, so it benefits from early attention.
Each of these activities belongs within the modernisation budget from the outset. Treating assurance as a contingency to be drawn upon when problems emerge tends to cost more, since the same work is then carried out under greater time pressure and with less opportunity to influence the design.
Assurance as a foundation for modernisation
Test debt shapes the cost, pace and risk of modernisation as surely as architectural and data debt do, and it deserves the same attention in the business case. Organisations that measure it before approving a route can choose more confidently between options, set timelines that hold and release change with far greater assurance.
At Audacia, we begin modernisation engagements by assessing assurance alongside architecture and data. That assessment establishes which behaviours are known and tested, how quickly change can be verified and where test debt is likely to affect the chosen route. We have found that this early investment consistently repays itself, giving sponsors a clearer view of risk and giving delivery teams a firmer foundation on which to build.


