# Evergreen engineering beyond the context window

> A live engineering case linking sustained work, accountable change and continuous renewal beyond individual model contexts.

Paper: Empirical follow-on — 16 September 2026
Published: 2026-09-17
Canonical: https://harten.io/papers/evergreen-engineering-beyond-the-context-window/

Cover illustration: [View the title image](https://harten.io/media/papers/covers/v1/evergreen-engineering-beyond-the-context-window.jpg?v=e79c979c74d7). Illustrative cover artwork, not an application screenshot or a record of measured results.

## Sustained work, accountable change and continuous renewal

By Lavan Nallainathan · Harten Technologies

Case observed 16 September 2026. Published 17 September 2026.

## Abstract

This paper reports live engineering and testing at Harten in which three concurrent workstreams continued as task contexts, implementation candidates and operating decisions changed. The records include continuation after context compaction, review repeated when relevant source state changed, preservation of unrelated work, unsuccessful qualification attempts and a measured improvement in a deployed interface. They also record two workers completing their tasks and leaving, followed by explicit authorisation for a bounded continuation by the remaining worker. These are empirical observations under the reported operating conditions, not merely proposed capabilities. The interpretation is that an engineering process can outlast its individual participants and reasoning contexts when its relevant knowledge, evidence and authority remain available. We relate the case to existing research and consider its implications for continuous renewal and sensitive, high-consequence applications. Independent replication and workload-specific security assurance remain distinct from the results reported here.

## 1. A process that outlasts its participants

In [Six hours of work, and a decision that remained mine](https://harten.io/papers/six-hours-of-work-and-a-decision-that-remained-mine/), I described roughly six hours of engineering execution with my interventions concentrated on review and direction. The work still ended with an unresolved quality decision. Substantial execution had taken place without the authority to accept an inadequate result passing to the agent.[1]

[Coordinating autonomous software agents under shared mutable state](https://harten.io/papers/coordinating-autonomous-software-agents-under-shared-mutable-state/) examined a different difficulty: several agents working against a changing system without invalidating one another’s work or the evidence supporting it.[2]

On 16 September, these questions came together in a longer operating day. Engineering continued from about 11 am into the late evening. That duration is my account of the working period, not a continuous instrumented record of unattended execution. I remained involved, and the evidence does not establish that all three workstreams were active throughout the whole period.

The records show real work continuing after context compaction, concurrent activity being coordinated, candidates being tested and corrected, and completed tasks being closed while other authorised work remained open. This is a participant case report based on contemporaneous engineering records and my observations.[3] Such investigation of a working system in its natural setting is an established empirical method in software-engineering research.[4]

The case concerns Harten’s use of AI-assisted engineering in the development of MaaS, its governed modernisation platform. It is not a claim that every capability described is native to one model or execution client. The reported behaviour belongs to the combined operating system around the work.

## 2. What the live work demonstrated

The most useful evidence is not that three sessions were busy. It is that progress and acceptance remained distinguishable while the work changed.

The interface work went through repeated qualification attempts. Several candidates improved measured performance but still missed the agreed acceptance threshold. Their improved results were reported without being relabelled as a pass. Independent review also found a cancellation defect, which was corrected and tested before the affected candidate proceeded.[3]

One later browser check reported the first status result appearing at 458 ms and all displayed status results arriving within 582 ms, compared with an earlier navigation observation of roughly 21.7 seconds. That is a reported improvement in a deployed interface. It is one browser-navigation comparison, including network and rendering, not an API percentile benchmark or an every-request guarantee. The underlying tests were performed during the engineering work; they were not independently rerun for this publication.[3]

The record subsequently reports the interface work deployed and its task-owned proposals resolved. I viewed the resulting interface, confirmed that it looked good and asked that task to stop. That is an operator assessment and a reported completion, not an additional quantitative benchmark.[3]

Unsuccessful outcomes remain part of the case. A candidate that failed qualification did not become acceptable because another component improved. A completed test did not stand in for review of the intended result. A later correction did not rewrite the outcome of an earlier attempt.

This matters because an engineering process can produce a persuasive history of activity while losing the distinction between what was attempted, what improved and what was accepted. The records here preserve those distinctions through a sequence of actual changes.[3]

## 3. Requirements can change without authority becoming ambiguous

Not every original condition remained unchanged. Two decisions concerned the meaning and cost of the service, not simply its implementation.

For one display, the agent asked whether a clearly timestamped snapshot up to 30 seconds old would be acceptable. I approved that boundary. The record distinguished that display allowance from the current information needed for commands and execution decisions. Elsewhere, I approved a running-cost trade-off to improve responsiveness.[3]

Those are authorised changes to operating requirements. They are not equivalent to an agent silently relaxing a threshold so that its candidate can pass.

The observation therefore supports a more useful statement than “the instructions never changed”. The system retained the distinction between working within an agreed condition and returning to the decision-maker when that condition needed to change.

That distinction also has to survive a context boundary. Recovering the first instruction is insufficient when a later decision has superseded it. Recovering only the latest instruction is insufficient when its scope was narrow. Permission to use older information for a display must not become permission to use it for every operational action.

My role remained consequential but bounded: select the objective, resolve material trade-offs and decide what further work to authorise. I was not required to relay every implementation or coordination transition.[3]

## 4. Coordination adapted when the work changed

The workers did not merely follow an unchanged queue. When one reviewed candidate required correction, another workstream had a ready dependency that was holding up its own authorised activity. The peers agreed a temporary change in the order of shared-source work, preserved the candidate awaiting correction and returned the opportunity to continue afterwards.[3]

That is evidence of coordination responding to the needs of the work. It does not quantify a time or cost saving. The narrower result is that an actual dependency was handled without every transition being scheduled by the human and without discarding the other workstream’s state.

Independent work continued where it did not conflict. When the relevant shared state changed, affected work was reviewed against the changed conditions rather than treating an earlier approval as indefinitely transferable. The contemporaneous records include preservation checks and explicit dispositions for superseded attempts.[3]

These observations extend the earlier two-workstream case. They do not establish that adding any number of workers will improve throughput. Accepted work, waiting, repeated review and intervention all matter to that question. More activity is not automatically more useful engineering.

## 5. Finishing a task is not finishing the programme

Later in the day, two workers completed their respective tasks and released their sessions in sequence. Their close notices retained the distinction between their own completed work and outstanding obligations belonging to the shared programme. The remaining worker acknowledged their departure without treating it as an instruction to close everything.[3]

This is a different result from keeping one agent running for longer. The participating set reduced while the work’s history and unresolved obligations remained identifiable.

The two departures occurred within the same operating case. They are not independent replications. They were orderly closures, not crash-recovery tests. The native close records and final deployed state were not independently reread for this manuscript; the supplied reports are the evidence for those particular operations.[3]

After the second closure, I asked whether the remaining worker should continue. It proposed diagnosing saved failures and preparing a correction, explicitly excluding a new experimental run. I approved that bounded proposal, and it acknowledged the same limit.[3]

The approval attached to the proposal, not to the word “continue” in isolation. Completed tasks remained completed. The saved failures remained failures. Preparing a correction did not authorise the next experiment or constitute evidence that the correction worked.

This gives continuity a practical meaning. The process needs to preserve what has finished, what remains owed and what is currently authorised. Retaining only “the work is complete” would stop the wrong activity. Retaining only “continue” could reopen finished work or carry an earlier permission into a different task.

## 6. Continuity is more than remembered text

A finite model context does not have to define the lifetime of an engineering process. In this case, the records show work continuing after context compaction while relevant state was available outside the immediate conversation.[3]

The important distinction is between retaining information and recovering its current meaning. An old approval may still be accurately recorded while no longer applying. A review may be correct for the candidate it examined while saying nothing about a later candidate. A failed result may be essential evidence for the next decision without becoming an accepted outcome.

The same issue applies to independent judgement. Preserving execution history is not enough when a reviewer or decision-maker has to reconstruct the intervening objectives, changes and unresolved questions. The evidence needed to judge current work must remain available without collapsing the distinction between execution and acceptance.

A recently retrieved record is not necessarily a recent observation. A summary is not the event itself. Recovery must preserve uncertainty and supersession, rather than turning a compressed account into a stronger claim than its sources support.

The interpretation of this case is that durable state and explicit authority support a continuing engineering process. The observed result belongs to the combined arrangement, including the models, execution environment, records and human decisions. The case does not isolate the causal contribution of each component.

## 7. The legacy-to-evergreen cycle

The consequence for my working day is a shorter interval between recognising a constraint and engineering a response. An implementation can be useful, expose a limitation and be reconsidered within days. Understanding the problem, reviewing a change and checking its consequences are increasingly part of the same ongoing process, rather than activities separated by a later recovery programme.

That comparison with months or years is my observation of a changing engineering cadence, not a measured industry baseline or a numerical acceleration claim. The evidence in this case supports rapid candidate evolution, review, integration and live improvement. It does not establish that every application can be modernised in days.[3]

Age alone does not make software legacy. A recently written component can become an inherited obligation when other parts of the system depend on its behaviour after requirements have moved on. An older component can remain suitable when its behaviour, dependencies and limitations are understood.

Evergreen, as used here, is the sustained ability to recognise and renew those obligations while preserving understanding and control. It does not require permanent churn. It does not mean there will never be a constraint or a need for replacement. It means the capacity to make justified change remains available as part of ordinary engineering.

For Harten, that capability is more important than any particular implementation’s age. The observed system enables useful work to continue without asking the founder or a single model context to reconstruct the entire history at every transition.

## 8. Why this matters for sensitive work

Sensitive, high-consequence applications make the distinction between capability and authority especially important. The ability to inspect information must not imply permission to alter it. The ability to propose a change must not imply permission to release it. Continued activity must not weaken the boundary protecting the workload.

Authentication, authorisation and accounting provide established security terms for identifying an actor, deciding what it may do and recording its activity. AAA describes these functions; it is not a security rating.[11] Identity, access management and audit information are also distinct concerns in the NCSC’s cloud security principles.[12]

Applied to persistent engineering, the requirement is that recovered knowledge remains subject to current access controls. Remembering a previous permission must not recreate permission that has expired or been revoked. Information derived from an application needs appropriate protection alongside the application itself.

The case demonstrates relevant engineering-control behaviour: bounded work, preserved evidence, explicit changes of authority and workers stopping when their tasks finish. Those results make sustained, accountable renewal a practical proposition for complex workloads. They do not replace assessment of the environment in which sensitive information will be processed.

Suitability for a particular workload must be established against its threats, data-handling requirements, deployment and assessment evidence. NIST’s control-assessment methodology likewise evaluates controls in context rather than treating an architectural description as sufficient proof.[13] The implication of this work is greater scope for controlled engineering, not an unrestricted security-accreditation claim.

## 9. Relationship to existing research

External memory, long-running agents and frequent software change are established research topics. Harten’s contribution here is the connected empirical account of sustained work, peer coordination, explicit human decisions and local task completion within a continuing engineering process.

Packer and colleagues’ *MemGPT: Towards LLMs as Operating Systems* studies management of information across immediate context and external memory. Its evaluations concern document analysis and multi-session conversation. It establishes a relevant precedent for separating the lifetime of a reasoning context from the information available to the wider system, rather than validating the particular engineering results reported here.[5]

Wu and colleagues’ *LongMemEval* distinguishes information extraction, cross-session reasoning, temporal reasoning, knowledge updates and abstention. These distinctions sharpen the recovery question: retrieving a previous decision and identifying the current decision are different tasks. Harten is not reporting an evaluation on that benchmark.[6]

Huang and colleagues’ *Persistent Recursive Worlds Enable Autonomous Software Evolution* is a close comparison. Their EvoX Genesis system centres continuity on accepted project state while individual agents remain finite-lived. The authors report software formation, continuation and redevelopment. The overlap with a process outliving its workers is substantive; this case does not claim priority or superiority over their approach.[7]

Anthropic’s *Effective harnesses for long-running agents* describes progress records, Git history, incremental work and completion checks across coding sessions. It is a primary engineering report rather than a peer-reviewed paper. Its relevance is the recovery of both completed and unfinished work across finite sessions.[8]

Kim and colleagues’ agent-scaling study evaluates 260 configurations across six benchmarks in the cited revision. Its task-dependent coordination results caution against treating additional agents as automatic throughput. Harten’s three-workstream case is another operating instance, not a general scaling curve.[9]

Shahin, Babar and Zhu’s systematic review of continuous integration, delivery and deployment identifies testing, visibility, design and infrastructure among the conditions affecting continuous practices. Frequent change itself is not new. The question here is how application and decision continuity are retained when part of the engineering workforce consists of model-driven processes with temporary contexts.[10]

These sources provide comparison and challenge as well as support. They do not endorse Harten or substitute for its evidence. This is a focused related-work selection, not an exhaustive literature review.

## 10. Findings and independent replication

The observed result is a live engineering process that continued across recorded context compaction, coordinated three workstreams, revised and tested implementation candidates, and retained distinctions between permission, review and acceptance. It includes a reported deployed improvement, unsuccessful qualifications, two orderly worker departures and a separately authorised continuation from saved evidence.[3]

The human role did not disappear. I selected objectives, authorised material trade-offs and decided what work should happen next. Peers handled several coordination exchanges without my relaying every step. That is a change in the division of work, not evidence of an unattended day.

Independent reproduction should report the operating conditions that could affect the result: model and client versions, workload, worker count, context boundaries, human interventions and acceptance criteria. The tests should distinguish accurate recovery of current decisions from simple recall of historical information, and accepted delivery from activity alone.

Rejected candidates, changed requirements, interrupted reviews and completed-task departures should remain visible in that evidence. Relevant measures include accepted work, lost or duplicated effort, coordination delay, recovery accuracy, cost and human intervention. Deliberate crashes, lost messages and restart are separate tests from the orderly closures observed here.

We invite others to reproduce the observed continuation and coordination under comparable, clearly stated conditions. This paper is not accompanied by a complete public reproduction package, and independent replication is not claimed. The contemporaneous source records remain private; that limits external auditability and is part of the evidence status, not a reason to describe the observed work as hypothetical.

The practical conclusion is that evergreen engineering does not require every worker to remain active indefinitely. It requires the understanding, evidence and authority needed for sound change to remain available as participants and contexts change. That is the capability exercised in this case, and the direction in which Harten’s evidence-first approach is developing.

## References

[1] Lavan Nallainathan, [Six hours of work, and a decision that remained mine](https://harten.io/papers/six-hours-of-work-and-a-decision-that-remained-mine/), 10 September 2026, with a 15 September addendum.

[2] Lavan Nallainathan, [Coordinating autonomous software agents under shared mutable state](https://harten.io/papers/coordinating-autonomous-software-agents-under-shared-mutable-state/), 15 September 2026.

[3] Harten, contemporaneous engineering records and operator observations, 16 September 2026. Private evidence supporting this participant case report. Includes reports of tests, review, integration, deployment, context continuation and session closure. The underlying results were not independently reproduced for this publication.

[4] Per Runeson and Martin Höst, [Guidelines for conducting and reporting case study research in software engineering](https://link.springer.com/article/10.1007/s10664-008-9102-8), *Empirical Software Engineering* 14, 131–164 (2009). Peer-reviewed methodology.

[5] Charles Packer and colleagues, [MemGPT: Towards LLMs as Operating Systems](https://arxiv.org/abs/2310.08560v2), arXiv:2310.08560v2, 12 February 2024. Research preprint; first submitted October 2023.

[6] Di Wu and colleagues, [LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory](https://arxiv.org/abs/2410.10813v2), ICLR 2025; arXiv version 2, 4 March 2025.

[7] Beichen Huang, Zhenyu Liang, Bowen Zheng and Ran Cheng, [Persistent Recursive Worlds Enable Autonomous Software Evolution](https://arxiv.org/abs/2608.10450v3), arXiv:2608.10450v3, 16 August 2026. Research preprint; results attributed to the authors.

[8] Justin Young, Anthropic, [Effective harnesses for long-running agents](https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents), 26 November 2025. Primary engineering report.

[9] Yubin Kim and colleagues, [Towards a Science of Scaling Agent Systems](https://arxiv.org/abs/2512.08296v3), arXiv:2512.08296v3, 8 April 2026. Versioned research preprint.

[10] Mojtaba Shahin, Muhammad Ali Babar and Liming Zhu, [Continuous Integration, Delivery and Deployment: A Systematic Review on Approaches, Tools, Challenges and Practices](https://doi.org/10.1109/ACCESS.2017.2685629), *IEEE Access* 5, 3909–3943 (2017). [Author manuscript](https://arxiv.org/abs/1703.07019). Peer-reviewed review.

[11] IETF, [Diameter Base Protocol, RFC 6733](https://www.rfc-editor.org/rfc/rfc6733.html), October 2012. Cited for authentication, authorisation and accounting terminology; no protocol conformance by Harten is asserted.

[12] UK National Cyber Security Centre, [The cloud security principles](https://www.ncsc.gov.uk/collection/cloud/the-cloud-security-principles). Accessed 17 September 2026. Guidance, not an assessment or endorsement of Harten.

[13] NIST, [SP 800-53A Revision 5: Assessing Security and Privacy Controls in Information Systems and Organizations](https://csrc.nist.gov/pubs/sp/800/53/a/r5/final), January 2022. Accessed 17 September 2026. Assessment methodology, not a Harten certification.

*This is a dated empirical follow-on within The Harten Papers, not a replacement for the earlier case reports. Related writing and the complete programme are available in [the Harten papers archive](https://harten.io/papers/).*
