New KPMG Agentic ERP, now available with KPMG in 140+ countries Read the announcement

Blog/Engineering

We gave software an immune system

Author
Yuzhong (WeiWei) LuoMember of Technical Staff
Published
September 21, 2026
Reading
7 min read

Light runs around 100 customer-facing apps on top of our public API. When a field in that API gets renamed, an app can keep asking for it by its old name. The request succeeds and the page loads, but the data never arrives. Where the customer should see a number, there is an empty space.

Nothing crashes. No alarm goes off. The first sign that something is wrong is often a customer reporting it.

That report starts a process that can take hours, sometimes days. Someone has to investigate, reproduce the problem, find the right engineer, work out a fix, test it, ship it, and tell the customer it is resolved.

Much of traditional monitoring assumes that broken software makes itself heard. Silent errors can slip past it entirely.

Two kinds of failure

Some failures are obvious: a page will not load, a database stops responding, or an automated test fails. They give us a clear signal to investigate, and our checks often catch them before a change ships.

Other failures are harder to see. An app requests data in a way that no longer matches the API. A pricing table quietly goes out of date. The software keeps running, and people keep trusting it, even as it returns incomplete information or corrupts data.

Two kinds of error. A loud error is caught the moment the change is proposed: the check runs, the change is blocked, and nothing reaches a customer. A silent error passes every check: the field is renamed, the request still succeeds, and a blank number sits on the customer's screen for days while nothing goes red.
The loud one is a bug. The silent one is a bug nobody is looking for.

These failures made us ask a broader question: could software notice when something was wrong, understand what had changed, and begin repairing itself?

That ability to respond and improve is part of what we call Organic Software. To build it, we had to look beyond crashes.

Looking for what monitoring misses

Our pipeline has four stages: detect, shape, file, and fix. At every stage, it can stop if proceeding would be unsafe or unnecessary.

The 4 stages a problem passes through: Detect, where scheduled checks find a silent failure; Shape, where the detector turns it into a structured Finding carrying its evidence and the limits of the fix; File, a single API call to the agent; and Fix, where the agent writes the change and a person approves it.
Detect, shape, file, fix. Every stage is allowed to stop.

Detection starts with scheduled checks for problems that ordinary monitoring can miss. We look for differences between the live API and its published specification, stale pricing tables, and apps whose descriptions no longer reflect what they do. We also check our test data, watch for errors as they happen, and verify that the scheduled checks are running at all.

That last check came from experience. On August 26, an incident in GitHub Actions, the service that runs our checks, prevented every scheduled check from running. Nothing turned red and nobody was notified. To the system, a check that never ran looked exactly like one that had passed.

We added a check for the checks. It was a small change, but it exposed a larger requirement: a system that repairs itself also needs a way to recognise when it has stopped looking.

Turning a problem into a safe task

Detecting a problem is only the beginning. The harder part is describing it clearly enough for another system to act on it safely.

Every detector produces a structured report called a Finding. It describes what is wrong, the evidence behind it, the affected files, the proposed replacement, and our confidence in the diagnosis. It also defines the limits of the repair.

Once the Finding is ready, filing it is an API call to Sirius, our coding agent, which picks up the task and works on the fix.

Those limits matter as much as the diagnosis. If resolving the problem requires changes outside the permitted scope, the agent must recognise that and stop. A narrow repair should never become an open-ended attempt to make the task look complete.

A Finding, the structured report a detector produces. It records what is wrong, the evidence, the replacement field name, which files are affected and how confident the detector is. Below a divider it sets the limits of the fix: what the fix may do, and what it may not.
A real Finding. The final row defines what the agent is not allowed to change.

An incorrect fix can be more dangerous than an unresolved finding. Once the system reports that a problem is handled, we stop looking for it. If the repair is wrong, the original bug remains, now hidden behind the assumption that it has been fixed.

Teaching the agent when to stop

Preventing that meant giving the agent a clear way to report two things: there is nothing to change, or the change cannot be made safely. Two early cases showed us why both outcomes matter.

Case 1

Nothing to change

In the first, the agent was asked to rename a field from required to isRequired. No code used the field, so there was nothing to update. But the instructions gave the agent no way to report that. It edited our local copy of the API specification instead, producing a change that looked plausible without addressing any actual problem.

It had no way to report that no change was needed

Case 2

Already fixed

In the second, an app was flagged for using primary instead of isPrimary, even though it had been fixed two days earlier. Again, the task offered no path for reporting that the work was already done. The agent changed the stand-in test data instead, breaking a protective check.

It tried to repeat work that was already done

Both cases exposed a gap in our instructions. We had asked the agent to make a change without accounting for situations where no change was needed.

We revised the briefs to make those outcomes explicit. If nothing uses the field, the agent can report that no change is necessary. If a repair exceeds the safe scope of the task, it can stop and explain why.

Knowing when to leave the code alone is an essential part of maintaining it. The system needs to recognise a successful outcome even when that outcome produces no code.

Keeping the system quieter than the bugs

Running the pipeline every day brought another lesson: automation only helps if it reduces the work and noise around a problem.

We cap the number of tasks a single run can create. If a detector suddenly reports 200 failures, generating 200 pull requests is unlikely to be useful. The detector itself may need attention, so the system should flag the situation for a human.

We also prevent repeated findings from creating duplicate work. A problem that appears on three consecutive nights should remain one task.

We learned this after a CI check ran with duplicate detection disabled.

Over a weekend, one problem generated 4 near-identical pull requests. The system followed its instructions exactly; those instructions simply failed to account for work already in progress.

Keeping people in control

It is tempting to take the idea of self-healing software straight to full autonomy. We are starting with explicit human approval.

Before the system can put forward a fix, it needs sign-off. Silence is never treated as consent, and every fix still passes through a person.

The work around that decision has changed. Instead of starting with a customer report and an open investigation, the reviewer can read the agent’s explanation, assess the proposed change, and either approve it or send it back with feedback.

A different experience of software

This began as a practical problem in our own codebase, but it points toward a different relationship with software.

Imagine something breaks and the software notices before you need to report it. It tells you what went wrong and that a repair is underway. Half an hour later, the problem is resolved.

The same accounts receivable summary with Outstanding showing 56,850 euros again. A note underneath reads: Outstanding is back. Broke at 09:12, noticed at 09:14, fixed at 09:41. No ticket was raised, because nobody had to raise one.
It broke at 09:12, was detected at 09:14, and was fixed by 09:41. There is no customer ticket because nobody had to raise one.

Today, we accept a lot of friction as part of using software. We file tickets, wait for engineers, read release notes, and learn workarounds for bugs that have existed so long they start to feel like features. We think software should take on more of that work itself.

Self-healing is our second step toward Organic Software. The first was adaptability, explored by my colleague in The First Step to Organic Software.

This pipeline is still early. It has detected real silent failures and produced real fixes. It has also made the difficult parts clearer: ambiguous findings, noisy detectors, outdated assumptions, unsafe repairs, and knowing when to stop. As the codebase grows, the system will need to extend its coverage and keep its understanding current.

That is the engineering work ahead of us. Each improvement brings us closer to software that can recognise a problem, respond to it, and recover, with less effort from the people who rely on it.

That is what we mean by giving software an immune system.

The apps this pipeline maintains run inside Light, the agentic ERP. Explore the App Store to see what Light already connects to, or book a demo.

Share All posts

Ready to build what's next?

Bring your ideas to life with Light. Explore what we can create together.

Book a demo