← Writing

The Pipeline Was the Attack

On 31 March 2026 a poisoned axios release went up on npm, and the machinery the industry built to keep itself patched delivered it to 895 repositories before anyone woke up. CISA's advice afterwards was to add seven days of latency. Microsoft's was to turn the bots off. Both are right, and the distance between them is the only decision left: security is a loop now, and the whole job is choosing what delay you set on it.

14th September 2026 · 13 min read

At 00:21 UTC on 31 March 2026, axios 1.14.1 went up on npm with a dependency nobody had asked for. Five minutes later Dependabot opened its first pull request to adopt it. npm’s automated scanners flagged the package as malicious at around the six minute mark. Removal took three hours. In the gap between those two facts the entire remediation apparatus of the JavaScript ecosystem ran exactly as designed and shipped a North Korean remote access trojan into production.

GitGuardian found at least 895 repositories that took a malicious version. Automation opened 154 of the pull requests and 95 were merged into main, 50 of them by a bot account with no human involved at any point. On jhipster/generator-jhipster the upgrade fired forty minutes after publication and merged sixteen minutes later, two hours before npm pulled the package.

Nobody in that chain made a mistake. Every one of those repositories had done the thing the industry has spent a decade asking for: keep your dependencies current, automate the boring part, do not let a known CVE sit in your tree for six months because a human forgot. The pipeline worked. The pipeline was the attack.

CISA’s alert of 20 April told everyone to set min-release-age=7 and install nothing that had been public for less than a week. Microsoft, the day after the incident, told everyone to pin exact versions and disable Dependabot and Renovate for critical packages outright. One agency said slow the loop down. The other said stop it.


The queue is not coming back

For two decades vulnerability management was a queue: something scanned, something scored, a ticket got raised, and a report went to a committee once a quarter. The whole structure rested on an assumption nobody wrote down: vulnerabilities arrive at a rate a human triage function can drain. That assumption is dead. NIST moved the National Vulnerability Database to a triage model on 15 April 2026, committing to enrich only an estimated 15 to 20 percent of incoming vulnerabilities. Volume went from 28,818 CVEs in 2023 to projections above 60,000 for 2026, and asking which function is actually in play cut 78 to 89 percent of SCA false positives. Nobody funded a doubling of the security team to meet it.

Sorting was the wrong problem. A queue drains at the speed of the humans attached to it, and no arrangement of humans drains 60,000 items a year. A loop drains at the speed of whatever verifies its output, and if that verifier is cheap the loop runs continuously and never has to be right about priority in the first place. The industry did not choose loops because loops are fashionable. It chose them because the other structure stopped working.


Cyber is where the verifier was already free

I have argued elsewhere that the agent is disposable and the harness is the asset, because reliability has left the weights and moved into the loop that checks the output and runs the model again when it fails. The objection is always cost: in most of the enterprise the check is a person reading a diff, which is why so much agentic work stalls.

Security remediation is the exception, for an unglamorous reason. The oracle is already written and already free. The exploit reproduces or it does not, the build is green or it is not, the vulnerable symbol appears in the call graph or it is absent, the CVE is in the SBOM or it has left. Nobody has to be persuaded and no committee has to agree, and the check costs a CPU cycle rather than an engineer’s afternoon.

DARPA priced it in public: the AI Cyber Challenge final put 63 synthetic vulnerabilities across 54 million lines of real open source in front of seven autonomous systems, which found 54 and patched 43, at $152 per task and 45 minutes per patch, every system open sourced afterwards.

That explains why the loop is possible, not why it landed here. A failing test is machine checkable, so is a performance regression, so is a stale dependency, and none of those has an industry of autonomous agents pointed at it. What security has that the rest of software engineering does not is somebody obliged to act on the finding, with a clock attached and a signature at the end. Most software has always been riddled with bugs and the industry has always been comfortable with that. Security is the part where the comfort became non-compliant. Cheap verification makes the loop possible and a compliance framework makes it fundable, and cyber is currently the only part of software where both hold at once.

The EU Cyber Resilience Act’s reporting obligations commence on 11 September 2026, and it is the first instrument here that binds whoever shipped the software rather than whoever runs it. An actively exploited vulnerability now requires an early warning to your national CSIRT within 24 hours of becoming aware, full notification within 72, and a final report within fourteen days of a fix being available, against fines up to fifteen million euro or 2.5 percent of worldwide turnover. The practical deadline for knowing what is inside your product is September 2026, not December 2027.


The fix already exists

Most security work is not authorship at all. Open source is 60 to 80 percent of a modern application. Somewhere between 77 and 95 percent of the vulnerabilities in a codebase live in transitive dependencies rather than in code your organisation wrote, and Black Duck puts the average application at 581 findings. For almost all of them, somebody upstream fixed the problem before you knew you had it, published the fix, and moved on. The remediation is not a thing you write. It is a thing you adopt.

That collapses the action space to almost nothing, because for the overwhelming majority of findings the output of the loop is a version string, which is why the package manager rather than the scanner became the place remediation actually happens.

The failure mode is the one most organisations are living in right now. A loop wired straight from scanner to pull request, with no gate on it, is not a loop. It is a hose. Most of what the scanner raises is not callable, so the machine generates churn at machine speed and the cost lands on review capacity, which was already the scarce input. Reachability and known exploitation status belong in front of the pull request, not on a dashboard behind it.


Five years running the manual version

I ran this by hand for five years, under the constraint that makes patching genuinely hard, which is a calendar somebody else owns. The CDN we built at Optus Sport was a two-engineer public facing open source stack: Varnish and Nginx and HAProxy over the kernel, OpenSSL underneath, a Go orchestration application on top. Every one of those components produced a steady arrival of critical CVEs that had to be triaged, tested and rolled out without dropping a live match. It was OTT-only, no broadcast fallback. A dropped match is a service outage, not a degraded backup.

So the loop ran at two speeds. Off peak we already had a loop in 2021: triage, test, roll, verify, then do it again before the next window, and the models compress the same pass. During an English Premier League weekend, the FIFA Women’s World Cup 2023, and EURO 2024 it froze hard. The only thing that made the freeze a decision rather than drift was that every deferred patch carried a tracked reason, an owner and a window it had to close in. Someone had to be able to answer, at a bad hour, what was deferred, why, until when, and what it left exposed. That was the whole control.

Five years without a CVE driven incident is the sentence people want from a run like that. There were incidents. Things arrived that had to be handled at speed and the schedule bent to them more than once. When you work near the bleeding edge, expect to find CVEs nobody else has discovered. What did not happen is that any of them turned into an outage. A loop does not stop things happening. It stops them becoming events, and the difference between those two sentences is the difference between having been lucky and having built something.

The clean version is the wrong lens anyway, because it is blind to the thing that would hurt you today. Nothing about axios 1.14.1 was a CVE. It was a valid signed version published by a compromised account and adopted by machinery running exactly as designed, and a spotless record against a scored list tells you nothing about whether you would have merged it. The parameter is the credential rather than the outcome: how fast the loop ran, how long it was allowed to stay open, and who owned that number.

The most useful thing I did in that whole run was not an improvement to the loop. It was migrating the TLS surface off OpenSSL and onto AWS-LC. OpenSSL was generating more urgent patch windows than everything else combined, and the 1 to 3 jump was too slow: we wanted the crypto on AVX on the EPYC cores, not offloaded to the NIC. The cheapest available fix was to stop having that dependency. When the loop is expensive, the move is not to run it faster. It is to shrink the surface it has to run over. It is still the only lever that reduces the number of findings rather than the cost of processing them.


The air gap is a dial, not a wall

An air gap gets argued as a wall and it does not behave like one. It behaves like a delay. The industry has just finished shipping that same control as a configuration line, with npm adding min-release-age in February 2026, Yarn and Bun shipping minimum age gates across late 2025, and Renovate having had the setting for years. Every major package manager now ships an air gap you set in days. The air gapped network is the same dial welded to a much larger number, and defended on principle rather than on arithmetic.

Where it works, it works completely, and axios is the proof. Any organisation installing from a curated internal mirror with a soak period never saw 1.14.1, ran no incident response, rotated no credentials, and did not need CISA’s alert of 20 April because by then it had been irrelevant to them for three weeks. That deserves saying plainly by someone who spends most of his time arguing for the opposite posture. It was the only control in the entire incident that worked without a single person doing anything.

Where it fails is the same property pointed at the other threat. CrowdStrike’s 2026 numbers put mean time to exploitation at negative seven days, with 42 percent of exploited vulnerabilities under attack before disclosure, while median time to remediate a critical sits at 43 days and is getting worse. A seven day cooldown costs you a week against those clocks and buys almost everything. A quarterly transfer window costs ninety days, and for Australian operators it makes the Essential Eight’s 48 hour expectation arithmetically unreachable on precisely the systems the isolation was built to protect. That gap is almost always closed by writing an exception rather than by changing either the control or the framework.

The deeper problem is that the air gap severs the wrong half of the loop. Take the network away from an agent working on remediation and it can still decide. What it cannot do is fetch the artefact. The fix already exists, it is two hundred milliseconds away on a network you have deliberately chosen not to have, and no amount of model capability closes that distance.

So split the loop at the boundary rather than break it. Resolve what is required outside, then bring approved artefacts back in through the transfer procedure you already have. That demands a real internal mirror with provenance attached, and in most air gapped estates I have seen, what exists instead is a file share holding tarballs somebody copied in during a project that finished years ago. Fix the mirror before buying anything else, because every other control here is downstream of it.

The unit on that mirror has to be a container. The tooling is already there, open source and commercial: one image a person can audit, not a snowflake per box. The cooldown fires without a person. Owning the number still needs one: someone who sets the delay, owns the override, and signs what crosses.


What the dial does not buy

Set it deliberately and write the number down. The evidence for a short cooldown is unusually strong, because this class of attack gets caught fast: Socket’s average time to detection across a recent multi registry campaign was five minutes and fifty six seconds, and Datadog puts twelve hours as sufficient to have blocked axios outright. Three to seven days costs almost nothing and filters nearly every registry compromise of that shape.

Write it down so it is a decision with an owner rather than a default nobody chose, and build the override before you need it, because a cooldown with no exception path is a control you will disable in a panic at the worst possible moment.

Then be honest about what it does not buy. xz-utils 5.6.0 shipped on 24 February 2024 and the backdoor was not found until 28 March, four weeks of dwell in a public release, caught in the end because Andres Freund was irritated that his SSH logins had gone from 100 milliseconds to 500. Any cooldown anyone would actually run in production would have passed it straight through, and it would have scored close to nothing on every prioritisation model in use at the time. A cooldown is a bet that somebody else finds the problem before your window expires, and it only pays against attackers in a hurry. It buys you the fast burn account takeover, which is most of what is happening, and nothing at all against the patient adversary, which is the one that would actually end you.

The far end of the loop broke this year and nobody has a setting for that one. Anthropic’s Project Glasswing produced more than 10,000 high or critical severity vulnerabilities across eleven partner organisations in its first month, independently assessed at a 90.6 percent true positive rate, and of 530 disclosed bugs, 75 had been patched. curl closed its bug bounty in January 2026 and paused vulnerability intake altogether in July to protect the people doing the reading. A loop runs at the speed of its slowest verifier, and at the end of a disclosure that verifier is a person who did not ask for any of this.


The fix already exists. For most of what your scanner will raise this quarter it existed before you scanned, written by somebody upstream who has moved on and will never know your company’s name. Everything expensive about security work is the distance between where that fix is and where you need it, measured in registries you do not mirror, windows you cannot open, approvals you cannot get, and dependencies three layers deep that nobody has looked at since the project shipped.

A loop is the only structure that closes distance repeatedly, and what you set is how fast it runs. There is no setting safe in both directions at once: fast enough to beat the exploit is fast enough to carry the poison, and slow enough to catch the poison is slow enough to lose the race. Anyone selling you one number for that dial has not read the incident.

So pick it yourself, know which threat you chose against, and write down why. Then answer the harder version, the one the axios post mortems mostly skipped: what did your automation merge last quarter, which account approved it, and what could that account reach. Fifty of those pull requests merged a trojan with no human in the room, and every one of those organisations answered that question from a standing start, three weeks after it mattered.

Connect

Based in South East Queensland. Open to permanent and contract roles.

Tell me what you run. Email opens and gets read, and warm introductions are welcome.