Light runs around 100 customer-facing apps on top of our public API. When a field in that API gets renamed, an app can keep asking for it by its old name. The request succeeds and the page loads, but the data never arrives. Where the customer should see a number, there is an empty space.
Nothing crashes. No alarm goes off. The first sign that something is wrong is often a customer reporting it.
That report starts a process that can take hours, sometimes days. Someone has to investigate, reproduce the problem, find the right engineer, work out a fix, test it, ship it, and tell the customer it is resolved.
Much of traditional monitoring assumes that broken software makes itself heard. Silent errors can slip past it entirely.
Two kinds of failure
Some failures are obvious: a page will not load, a database stops responding, or an automated test fails. They give us a clear signal to investigate, and our checks often catch them before a change ships.
Other failures are harder to see. An app requests data in a way that no longer matches the API. A pricing table quietly goes out of date. The software keeps running, and people keep trusting it, even as it returns incomplete information or corrupts data.
These failures made us ask a broader question: could software notice when something was wrong, understand what had changed, and begin repairing itself?
That ability to respond and improve is part of what we call Organic Software. To build it, we had to look beyond crashes.
Looking for what monitoring misses
Our pipeline has four stages: detect, shape, file, and fix. At every stage, it can stop if proceeding would be unsafe or unnecessary.
Detection starts with scheduled checks for problems that ordinary monitoring can miss. We look for differences between the live API and its published specification, stale pricing tables, and apps whose descriptions no longer reflect what they do. We also check our test data, watch for errors as they happen, and verify that the scheduled checks are running at all.
That last check came from experience. On August 26, an incident in GitHub Actions, the service that runs our checks, prevented every scheduled check from running. Nothing turned red and nobody was notified. To the system, a check that never ran looked exactly like one that had passed.
We added a check for the checks. It was a small change, but it exposed a larger requirement: a system that repairs itself also needs a way to recognise when it has stopped looking.
Turning a problem into a safe task
Detecting a problem is only the beginning. The harder part is describing it clearly enough for another system to act on it safely.
Every detector produces a structured report called a Finding. It describes what is wrong, the evidence behind it, the affected files, the proposed replacement, and our confidence in the diagnosis. It also defines the limits of the repair.
Once the Finding is ready, filing it is an API call to Sirius, our coding agent, which picks up the task and works on the fix.
Those limits matter as much as the diagnosis. If resolving the problem requires changes outside the permitted scope, the agent must recognise that and stop. A narrow repair should never become an open-ended attempt to make the task look complete.
An incorrect fix can be more dangerous than an unresolved finding. Once the system reports that a problem is handled, we stop looking for it. If the repair is wrong, the original bug remains, now hidden behind the assumption that it has been fixed.
Teaching the agent when to stop
Preventing that meant giving the agent a clear way to report two things: there is nothing to change, or the change cannot be made safely. Two early cases showed us why both outcomes matter.
Case 1
Nothing to change
In the first, the agent was asked to rename a field from required to isRequired. No code used the field, so there was nothing to update. But the instructions gave the agent no way to report that. It edited our local copy of the API specification instead, producing a change that looked plausible without addressing any actual problem.
It had no way to report that no change was needed
Case 2
Already fixed
In the second, an app was flagged for using primary instead of isPrimary, even though it had been fixed two days earlier. Again, the task offered no path for reporting that the work was already done. The agent changed the stand-in test data instead, breaking a protective check.
It tried to repeat work that was already done
Both cases exposed a gap in our instructions. We had asked the agent to make a change without accounting for situations where no change was needed.
We revised the briefs to make those outcomes explicit. If nothing uses the field, the agent can report that no change is necessary. If a repair exceeds the safe scope of the task, it can stop and explain why.
Knowing when to leave the code alone is an essential part of maintaining it. The system needs to recognise a successful outcome even when that outcome produces no code.
Keeping the system quieter than the bugs
Running the pipeline every day brought another lesson: automation only helps if it reduces the work and noise around a problem.
We cap the number of tasks a single run can create. If a detector suddenly reports 200 failures, generating 200 pull requests is unlikely to be useful. The detector itself may need attention, so the system should flag the situation for a human.
We also prevent repeated findings from creating duplicate work. A problem that appears on three consecutive nights should remain one task.
We learned this after a CI check ran with duplicate detection disabled.
Over a weekend, one problem generated 4 near-identical pull requests. The system followed its instructions exactly; those instructions simply failed to account for work already in progress.
Keeping people in control
It is tempting to take the idea of self-healing software straight to full autonomy. We are starting with explicit human approval.
Before the system can put forward a fix, it needs sign-off. Silence is never treated as consent, and every fix still passes through a person.
The work around that decision has changed. Instead of starting with a customer report and an open investigation, the reviewer can read the agent’s explanation, assess the proposed change, and either approve it or send it back with feedback.
A different experience of software
This began as a practical problem in our own codebase, but it points toward a different relationship with software.
Imagine something breaks and the software notices before you need to report it. It tells you what went wrong and that a repair is underway. Half an hour later, the problem is resolved.
Today, we accept a lot of friction as part of using software. We file tickets, wait for engineers, read release notes, and learn workarounds for bugs that have existed so long they start to feel like features. We think software should take on more of that work itself.
Self-healing is our second step toward Organic Software. The first was adaptability, explored by my colleague in The First Step to Organic Software.
This pipeline is still early. It has detected real silent failures and produced real fixes. It has also made the difficult parts clearer: ambiguous findings, noisy detectors, outdated assumptions, unsafe repairs, and knowing when to stop. As the codebase grows, the system will need to extend its coverage and keep its understanding current.
That is the engineering work ahead of us. Each improvement brings us closer to software that can recognise a problem, respond to it, and recover, with less effort from the people who rely on it.
That is what we mean by giving software an immune system.
The apps this pipeline maintains run inside Light, the agentic ERP. Explore the App Store to see what Light already connects to, or book a demo.


