Harten Research — Research 01
Recursive Application Denoising
An evidence-directed approach to reconstructing complex software systems, compressing application discovery and directing human attention to the uncertainties that still require judgement.

Evidence-directed reconstruction of complex software systems
Before an organisation can modernise a complex application, it first has to understand it.
That sounds obvious. In practice, it is where a surprising amount of time disappears.
Legacy applications rarely arrive with a current, complete description of how they work. Code has accumulated over years. Business rules are distributed across application logic, databases, configuration, integrations and operational processes. People leave. Documentation drifts. Exceptions become normal behaviour. Dependencies remain because something somewhere still expects them.
So modernisation starts with reconstruction.
Teams inspect the estate, speak to subject matter experts, trace dependencies, compare documentation with implementation and try to establish what the application does, where risk sits and what can change safely.
Application discovery is still commonly planned in weeks rather than hours or days. The exact duration varies significantly with scope, estate quality, access to subject matter experts and the depth of understanding required. The important point is not a single industry average. It is that substantial expert effort can be consumed assembling the initial account of a system before those people can exercise judgement over it.
That matters because modernisation programmes repeatedly get into difficulty before the replacement code is the main problem.
In February 2026, the UK Public Accounts Committee reported on NS&I's transformation programme. Among the issues it identified were a weak understanding of the highly integrated system being replaced, contracts awarded before dependencies were sufficiently understood, the absence of an integrated plan and capability gaps.
The National Audit Office has made the broader point that digital transformation is not greenfield delivery. New technology has to fit existing systems, data, processes, contracts and organisational constraints. The US Government Accountability Office has similarly highlighted incomplete modernisation planning for critical legacy systems.
These are different situations, but there is a recurring risk underneath them: consequential decisions are made while understanding of the existing system remains incomplete.
A target architecture can be coherent and still be based on an incomplete account of current behaviour. A migration can pass its new tests and still lose behaviour nobody thought to specify. A replacement can go live while the old system remains difficult to retire because an unexpected process still depends on it.
Modernisation therefore has two different problems.
The first is transformation: how to build and migrate the new system.
The second comes before it: how to establish a sufficiently reliable account of the system that already exists.
Harten has been working on the second problem.
The discovery bottleneck
Application knowledge is distributed.
An architect may understand the overall platform but not every business exception. A developer may understand an implementation path but not why the rule exists. A product owner may understand the intended service but not what the production implementation actually does. Operations may understand recurring failure modes without knowing where the corresponding behaviour lives in source.
The conventional response is to bring these people together and reconstruct the system.
That work is necessary, but it can be slow, difficult to repeat and dependent on the continued availability of scarce expertise.
AI changes the economics of this work, but only if it is used carefully.
A language model can analyse code and produce convincing documentation quickly. A convincing explanation is not the same thing as a defensible application model. If the evidence is incomplete, fluency can hide the gap rather than resolve it.
The objective therefore cannot be to ask a model to understand a repository and accept the answer.
The objective is to build an application-level account that remains connected to evidence, can be challenged, and makes uncertainty visible.
That is the problem Recursive Application Denoising is intended to address.
What Recursive Application Denoising means
Recursive Application Denoising is Harten's approach to progressively reconstructing a complex application from fragmented technical evidence.
The word denoising is an analogy. It is not a claim that the system is a diffusion language model or a new neural architecture.
A legacy application begins, from the analyst's point of view, as a noisy object. There are many local facts but no single reliable account of the whole. Concepts can be duplicated. Relationships can be missing. Technical components can be visible without their business meaning. Evidence can conflict. Some questions cannot be answered from source alone.
Harten builds and progressively refines an evidence-linked model of that application. Rather than treating generated documentation as truth, the approach preserves provenance, tests emerging interpretations against retained evidence, exposes unresolved uncertainty and directs human review towards questions that remain material.
The important distinction is that the output is not intended to be an oracle.
It is an evidence-backed application understanding that can be inspected and challenged.
Where the available evidence does not justify a conclusion, the question should remain visible rather than being filled with a plausible answer.
The detailed orchestration, evidence-selection, refinement and validation mechanisms are proprietary Harten implementation know-how and are intentionally outside the scope of this public paper.
A three-day reconstruction
We exercised the approach against the public `DEFRA/prsd-iws` repository, an approximately 292,000-line legacy .NET application supporting International Waste Shipments.
This was a Harten-controlled analysis of a public repository. It was not commissioned by DEFRA, is not a customer case study and should not be read as DEFRA validation or endorsement of the output.
The application was analysed over three days.
The exercised environment was constrained by available Azure compute quota. The workload could not be parallelised to the degree the architecture is intended to support. We therefore treat three days as the observed elapsed result of this run, not as a theoretical lower bound or an unmeasured scaling claim.
The reconstruction produced an evidence-linked application model and a reviewable documentation set. It retained unresolved uncertainty rather than presenting the output as complete or infallible.
That distinction matters more than a claim of perfect coverage.
The result we care about is that a substantial legacy application could be moved from fragmented implementation evidence to a structured account of the system in days, with specific questions surfaced for human review.
Human review is not a concession
There is a tendency in AI discussions to treat human review as evidence that automation has failed.
That is the wrong frame for modernisation.
The question is not whether a person is ever needed. They are.
The question is what we are asking that person to do.
In a conventional discovery process, an experienced engineer, architect or domain expert may spend days or weeks helping reconstruct the system from first principles.
An evidence-first process changes the starting point.
Instead of asking somebody to explain an application from a blank page, the review can begin with a reconstructed account of its behaviour, supporting evidence and the areas where that account remains uncertain.
That is still human judgement.
It is a more concentrated use of it.
There are also things source code cannot establish. It may not tell us which apparently unused integration is still contractually required. It may not prove the runtime topology. It may not contain a policy decision made outside the application. It cannot decide whether an organisation wants to preserve a behaviour in the target system simply because that behaviour exists today.
Those are boundaries, not defects in evidence-first discovery.
The aim is to reach those boundaries sooner and with the surrounding technical picture already assembled.
Does this change the economics of discovery?
The three-day result is not a universal benchmark. Applications vary enormously by language, architecture, repository quality, runtime behaviour and available evidence.
But the shape of the work is changing.
Understanding a large legacy system has traditionally been treated as predominantly manual work that scales with the availability of scarce experts.
The relevant question now is how much of the reconstruction can be performed computationally, how much supporting evidence can remain attached to the resulting understanding, and how effectively the remaining uncertainty can be handed to people qualified to decide what it means.
That is where Harten's proposition sits.
We are not claiming that a model makes an organisation infallible.
We are not claiming that an application is completely understood because documentation has been generated.
We are not claiming that three days of analysis is three days to modernise an application.
And we are not claiming that architects, developers or subject matter experts disappear from the work.
The claim is more practical.
A substantial part of application discovery can move from manual reconstruction towards evidence-directed computation, so that human expertise is spent reviewing uncertainty and making consequential decisions rather than rebuilding the entire system model by hand.
If that holds across more systems, it changes the economics of modernisation before transformation even begins.
Why this matters for modernisation programmes
Modernisation programmes do not fail simply because engineers cannot produce new code quickly enough.
They get into difficulty when uncertainty is discovered too late.
An unresolved dependency found during discovery is a review item.
The same dependency found during migration is rework.
The same dependency found after cutover can be an incident.
Better discovery cannot eliminate programme failure. It cannot create executive sponsorship, settle business priorities, fund the programme or resolve organisational disagreement.
But it can change when technical uncertainty becomes visible.
The earlier it is made explicit, the more choices the organisation still has.
That is why Harten treats application understanding as something worth retaining. The resulting evidence and specification should be available to support later planning, implementation and verification rather than disappearing when an assessment ends.
What remains to prove
One application is not a universal benchmark.
The method inherits the limits of its inputs. Behaviour absent from the available evidence cannot be proven by better reasoning. Traceability can show why an interpretation was made; it does not make every interpretation correct.
The three-day result was also recorded under constrained compute capacity. Greater parallel capacity is expected to affect elapsed time, but that remains an inference until measured on comparable workloads.
The standard Harten intends to work against is therefore practical: whether repeated application engagements reduce the elapsed time and human effort required to reach a useful reviewable baseline, while retaining enough evidence for experts to identify mistakes, omissions and unresolved questions.
The underlying change
For a long time, organisations have treated discovery as something people have to do before the real modernisation can start.
I think that boundary is changing.
Application understanding can itself become an engineered capability.
The result is not an oracle.
It is a better starting point for a difficult decision.
And in complex modernisation, a better starting point can be worth a great deal.
References
- UK Parliament, Committee of Public Accounts, NS&I's transformation programme, 13 February 2026.
- National Audit Office, Six reasons why digital transformation is still a problem for government.
- U.S. Government Accountability Office, Information Technology: Agencies Need to Plan for Modernizing Critical Decades-Old Legacy Systems, 17 July 2025.
- Source analysed by Harten: DEFRA/prsd-iws.
This is a standalone Harten research paper. The complete Harten papers archive is available at harten.io/papers.
Where this becomes operational
Apply the thinking to a real application.
If the problem described here exists in one of your applications, Harten can establish the current evidence, unresolved uncertainty and the basis for the next decision.