AI Security
Local Prevention Is Not the Hard Part
Modern EDR already decides on the sensor, so moving the decision to the endpoint is not the interesting problem. The interesting problem is what a fleet of local deciders still cannot do without re-centralizing. Version 2.0 of the Endpoint Mesh is a correction, and the corrections are more useful than the original claims.
August 14, 2026 · 18 min read

There is a version of the endpoint security argument that goes like this. Detection ships telemetry to a central place, the central place thinks about it, a verdict comes back, and by the time it lands the process has already spawned its children and written to disk. So the decision has to move to the endpoint.
The first half of that is true. The conclusion is ten years late.
Commercial EDR already prevents on the sensor and has for years. CrowdStrike Falcon Prevent, Microsoft Defender's block and attack disruption behavior, and SentinelOne's on agent Storyline all reach a verdict and act locally, including when the cloud is unreachable. For those products the cloud is hunt, intelligence, and after action work. It is not in the prevention path.
So "move the decision to the endpoint" is not an insight. It is the installed base.
The Endpoint Mesh was published on July 28 as a defensive publication, released as prior art so the mechanisms could not be fenced off. Version 2.0 is now out, and it is a correction. Eleven load-bearing claims from version 1 did not survive a proper review, and the framing above was one of them. Every withdrawn claim is quoted and explained in ERRATA.md.
The corrections turned out to be more useful than the original claims, so this article is about those. What is left after you grant modern EDR its due is narrower, harder, and more interesting than what version 1 was arguing.
What is actually left
Grant that the endpoint can already decide. Three problems survive.
Telemetry the center never receives. SIEM ingest is priced by volume, so every real deployment has a quiet tier of what does not get sent. Verbose process creation, full command lines, DNS, module loads. Those are precisely the signals that distinguish a novel attack from a noisy Tuesday, and they are the first things filtered at the agent or dropped at the pipeline for cost. A center cannot decide on data it was never sent. This is the strongest argument in the whole document and it is an argument about where the decision happens, not about what makes it. Rules and YARA get the same benefit as a model does.
The first hop. A host that just moved laterally knows something its destination does not, and the fastest path between them is the edge they just created, not a round trip through a console.
An action set an operator can sign off on. The constraint on autonomy was never detection quality. It is whether the action has an inverse that somebody trusts at three in the morning.
Notice that none of those is solved by making the endpoint smarter. They are solved, if at all, by what a fleet of endpoints is allowed to say to each other and allowed to do without asking.
The question worth asking
If a single endpoint can already decide, the useful question is what a fleet of local deciders still cannot do without re-centralizing.
Two things.
It cannot see the host next to it. Lateral movement is by definition a relationship between hosts, and the relationship exists at the two endpoints of that edge before it exists in any console.
It cannot be trusted with an unbounded action set. A fleet of confident local deciders that can do anything is an outage generator with good intentions.
Everything below is what happened when each of those was worked through properly, including the parts where the first version got it wrong.
Route on how fast it is moving, not on who is driving
Version 1 routed responses on actor class. If a classifier decided the attack was AI driven, it went down the fast autonomous path. If it decided the attack was "human, scripted, or throttled," it went to the SIEM for correlation with a person in the loop.
The same document, one section earlier, stated that machine agents decide faster than a human and slower than a fully scripted exploit.
Put those two sentences next to each other. Scripted attacks are the fastest thing on the network, and the router's fast lane excluded them. WannaCry, NotPetya, a ransomware encryptor working through a file share, a beacon with sleep set to zero. None of those is "an AI." All of them are machine speed, and they are the exact case the opening argument says cannot wait.
That is a fast lane that excludes the fastest traffic, and it survived several readings of the document because each section is individually reasonable and the contradiction only appears when you read section five against section six rather than in order.
There were structural problems underneath it too. The stated question was two class and free of malice, "is the actor a human or a machine, regardless of intent," but the classifier emitted three classes with malice baked into two of them. The training data filed existing malware corpora under "human driven," which teaches a model that automation is a person. And the tempo signal the whole thing rested on was already documented as beatable with timing jitter in an afternoon, which meant a beatable signal was wired straight to an actuator.
Version 2.0 routes on observed velocity, blast radius, and probability of malice. Things that are measured rather than inferred from a label: process spawn rate, file encryption rate, distinct destinations contacted per interval, time from first observation to privileged operation. Scripted machine attacks default to the fast path, because they are the fast case.
The actor label survives as a field on the ticket. It is genuinely useful for triage, for knowing what to expect next, and for writing the report. It is not a switch.
The nice property of this change is what it does to the jitter weakness. Under actor class routing, defeating the tempo signal moved an attacker into the slow path, which is the failure the document opened by describing. Under velocity routing, defeating the tempo signal costs you an accurate label on an alert and changes nothing about how the response is routed, because nobody asked the label.
You cannot un-kill a process
Version 1 stated a design rule that still holds up:
If an autonomous response cannot undo itself unattended, it is not ready to be autonomous.
Then it gave a representative action set that violated the rule one page later: kill the process and its descendants, isolate the network, snapshot the filesystem, invalidate the user's credential and Kerberos ticket. All committed atomically, all succeeding or none.
You cannot un-kill a process. You cannot locally un-purge a Kerberos ticket or re-enable an account without the directory, which means the network, which is the thing containment just cut. Isolation was the only action in that list that actually satisfied the rule.
The atomicity claim was the bigger problem. Windows and Linux have no transaction that spans the process table, the host firewall, a volume snapshot, and a remote KDC. Staging a kill without applying it is not an operating system primitive. Treating four subsystems as one operation in user space does not make them one operation. Version 1 also claimed the atomicity was a security property, that an attacker "cannot interfere with action three while action one is in progress," which a user space sequencer cannot promise to an attacker who already has privilege on the box.
And a live PID emitted by a model is stale before staging finishes. A model asked to produce PIDs, interface names, and SIDs will invent them.
The autonomous set in version 2.0 is three actions that have an inverse the host owns by itself:
- Freeze a process tree, using the cgroup freezer or thread suspend. It stops making progress, stays available for evidence, and thaws.
- Reversible local filter rules on the host firewall, with a time to live.
- A filesystem snapshot taken before any mutation.
Kill, credential invalidation, and account disable move to escalate only. Some can be issued as a request to an identity service able to re-issue, but that is a request, not a local action.
The engine is described as a compensating saga: stage, apply in order, verify each step against local state, compensate backwards on failure. Compensation can also fail. There is a real race in there, and the corrected version reports it rather than hiding it behind the word "atomic."
One more thing falls out of this. Version 1 listed "is the management channel still up" as a verification condition for a successful containment. If the reason for isolating is that the agent itself might be untrusted, the management channel being up is a failure condition, not a success condition. Version 2.0 splits isolation into a soft profile with a specified management allow list and a hard profile with no exceptions, and verification never depends on the management path.
Warn, then contain
Version 1 ordered the autonomous path as: respond locally, then tell the peers, then inform the manager. Isolation was described as preserving a management channel, so the ordering looked harmless.
It is not harmless. Isolation cuts the data plane. The entire peer mechanism depends on a compromised host sending a signal toward its neighbors, and after it isolates, it cannot send. The neighbors get silence, and silence reads closer to all clear than to alarm. A mesh built to warn its neighbors had an ordering bug that cut the warning.
Warn, then contain is now a protocol invariant. Stage the isolation, flush a signed alert to current neighbors on the still open path, then commit the isolation. The order is normative, not an optimization.
Three related rules came out of the same analysis. An isolated node may not author all clear messages, because a node that just contained itself is in no position to certify that anything is fine. All clears carry near zero weight regardless, because a false alarm costs an inspection and a false all clear costs the incident. And an attacker can sign a final message and then go dark, so a high reputation last word from a node that immediately goes silent waits for a second independent source before it counts.
Listening to a neighbor is not trusting a neighbor
Version 1 defined adjacency topologically. A peer is adjacent if you have communicated with it recently. A host that scanned you twelve seconds ago is an edge.
That is still the right graph for lateral movement. Who you just talked to is who could have moved, and an attacker pivoting from A to B creates exactly the edge that matters at exactly the moment it matters. That claim survives intact.
What does not survive is using the same edge for trust. The attacker chooses who they talk to. One sweep across a /16 makes the scanner adjacent to the entire fleet, and if adjacency conferred trust, one valid key plus one nmap run would let an attacker inject verdicts and lessons everywhere at once. Meanwhile a quiet local encryptor talks to nobody, is adjacent to nothing, and never triggers the mechanism at all.
Domain controllers, file servers, and software distribution hosts also become hubs by construction, because everyone already talks to them. Calling the design leaderless was true of the protocol and false of the graph.
Version 2.0 splits adjacency in two:
- Observation adjacency is recent communication. It earns a peer the right to be listened to, and it is a reason to go inspect your own inbound artifacts from that peer.
- Trust adjacency is policy: site, role, attested build. It is the only thing that earns the right to be learned from.
Scanning creates the first and never the second. High degree infrastructure nodes may report facts about sessions they actually terminated, but they may not contribute catalog entries and they may not autonomously isolate.
There is a second fix in how detection works at the receiving end. Version 1 required the compromised source to honestly self report, which is a strange thing to require of a compromised host. Version 2.0 runs inference on the receiver's own observations of that peer and treats the signed alert as an optional prior rather than as the event. If the source lies or goes quiet, the destination still detects. It just loses a hint.
And a peer message never fires containment on its own. A gossip layer with actuator rights is a fleet wide denial of service primitive with a valid signature on it.
A rule that only accepts what you already believe is not learning
Version 1's defense against poisoning was consensus before integration. A receiving endpoint would integrate a lesson only if its own local model would have classified the same way. Disagreement meant reject, and decrement the sender's reputation. The document called it the load bearing part.
It is backwards, and the reason is a single sentence: learning is an update on surprise.
A rule that integrates only what the receiver already agrees with accepts what everyone already knew and rejects genuine novelty. A new attack is precisely the thing your model would not have called the same way. The first host in a fleet to see something new could not teach anyone about it, and had its reputation docked for being first.
Worse, the poison it was built to stop walks straight through. Version 1 admitted in its own limitations section that plausible poison defeats consensus checking by construction, because consensus checking is exactly a plausibility test. Both facts were in the document. Neither was traced to the conclusion that the mechanism cannot do the job.
You cannot use one test as both your poison defense and your learning rule. It has to reject one thing and accept the other, and those are the same thing.
Version 2.0 stops calling this learning. It is indicator distribution with local corroboration:
- The object is a signed, versioned, size capped feature and indicator catalog entry. Not a LoRA adapter, not a prompt prefix, not a gradient. Version 1 described a sub 4 KB object that was somehow all three at once.
- Disagreement is candidate novelty. It goes to quarantine and is a reason to look harder, not to reject.
- Integration requires a second, independent local observation. The corroboration is evidentiary, not a model vote.
- Reputation moves only on proven falsehood, meaning the entry predicted a class and this host later confirmed the opposite outcome.
Version 1 also generated lessons from "the classifier fired and the response succeeded," with no human anywhere on that path. That is self certification: a model teaching the fleet what the model already believed, laundered through an action that only verified that the action landed.
What peer rumor cannot see
The mechanism that shares signals between neighbors was called correlation in version 1. It is not correlation, and the rename matters because it sets expectations correctly.
It is neighbor compromise rumor: a host that believes it is compromised telling the hosts it was recently talking to, quickly, in a small signed message. Score, class, origin, time, signature. Indicators become an optional pull for a receiver that decides it wants them.
Version 2.0 lists what rumor cannot see, which is most of what a SIEM is actually for. Kill chains split across hosts that never talk to each other. Identity centric abuse spread thin across many hosts. Anything in SaaS, the IdP, or email. Whether an artifact is first seen in the fleet. Population level beaconing. Low and slow campaigns measured in weeks.
Those need a center or they are not seen. Which is why the corrected framing is a local prevent path plus optional neighbor rumor sitting beside a SIEM that keeps every job it already earns: long horizon hunt, legal hold, compliance attestation, case management, and correlation across identity and email and SaaS. Version 1 claimed the four mechanisms performed every function a SIEM performs, in a document that also said the mesh complements the SIEM two sections later. Only one of those can be true, and it is the second one.
Honest status
Nothing in the four mechanisms is implemented. Not partially, not in a lab, not behind a flag.
Version 1 listed a project called Specter as the running implementation of the collection layer. Specter is a central Wazuh and Suricata dashboard that polls an indexer on a timer. Collection was also the one function the document said never needed the paper, because collection was always local. Citing it as implementation status was a category error, and version 2.0 says plainly that nothing supports the argument by way of running code.
There are also no performance numbers, because there are none worth reporting, and the fleet scale propagation target that appeared in version 1 has been deleted rather than relabeled. There was no protocol specified precisely enough to have a target for.
Version 2.0 is also explicit that it is not a complete specification. Membership, NAT traversal, a shared event schema, the packet path that carries the warning, a reputation formula, and a catalog codec are all missing. Field names are not a protocol. Version 1 said the mechanisms were described in enough detail to be built by anyone, and that claim is withdrawn.
One thing moved in the other direction. Version 1 placed key distribution and revocation outside its boundary while every mechanism assumed signed identity, which is the lock and the door. Version 2.0 brings identity inside: reuse the agent certificate, mTLS, or the enrollment channel the fleet already has, keep keys short lived, and treat revocation as a first class object that does not have to be independently corroborated before it is honored.
Why the correction is published instead of applied quietly
Version 1 shipped with a correction policy stating that a wrong claim would be fixed in a versioned revision with the change noted, rather than silently edited.
So version 1.3.2 stays permanently citable at its own DOI, mistakes included. Version 2.0 sits alongside it, and the concept DOI resolves to whichever is current. The original article is still up, uncorrected, with a notice at the top pointing here.
That is the whole difference between a disclosure and a marketing page. A marketing page gets edited. A dated record that quietly rewrites itself is not a record, and prior art that cannot be trusted to be the text it claims to be is worth less than nothing.
Three checks worth running on your own architecture documents
The corrections above have a shape, and it generalizes past this framework.
Grep for the word "every." Every slogan that had to be withdrawn contained it. Every function a SIEM performs. Any signed conclusion. All four mechanisms with no modification. In each case a narrower and more defensible sentence elsewhere in the same document already contradicted the slogan, and the slogan was doing marketing work in a document that was not supposed to have any.
Read your sections against each other, not in order. The routing rule and the tempo analysis contradicted each other on the single most important design decision in the paper. Reading forward hides that, because forward is the order in which each section sounds reasonable on its own.
For every limitation you disclosed, ask what load it was carrying. Version 1 had an honest limitations section. Timing jitter beats the actor signal. Reputation is a mitigation, not immunity. No fleet scale numbers exist. All true, all volunteered.
What was missing was the second step. "The tempo signal is beatable" is a footnote if tempo is a label on a ticket. It is fatal if tempo selects the response architecture, which is what it did. "Consensus checking is a plausibility test" is a caveat until you notice it is the same test you rely on to accept new knowledge.
Disclosing a weakness is not the same as tracing where it lands. Most of version 2.0 is the result of tracing weaknesses that version 1 had already written down.
It is still unowned
Any project, open source or commercial, is free to implement all of it. No permission needed, no attribution owed, and that covers the version 1 forms of these mechanisms too, including the ones just withdrawn. The only ask is that anyone building from it reads the errata first and does not build the parts that came out.
The concern that motivated publishing this in the open is unchanged: endpoint autonomy should not become a fenced yard where a handful of platform vendors are permitted to ship it and the open source security projects that most organizations actually run get locked out.
A smaller, correct framework is more likely to get built than a larger one that asks the reader to believe six things that are not true.
Version 2.0 of the full disclosure, the errata listing every withdrawn claim, and the patent non-assertion are at github.com/CarbeneAI/endpoint-mesh, archived at DOI 10.5281/zenodo.21716611, which always resolves to the current version. Version 1.3.2 remains permanently citable at 10.5281/zenodo.21716656 if you want the text as it originally shipped. The original article is still up, uncorrected, with a notice pointing here.