Canaries catch tampered nodes in P2P LLM inference
New research tackles integrity in peer-to-peer Large Language Model inference, where any node can corrupt intermediate activations. The authors mix secret canary inputs into traffic and rank nodes by activation drift. It nails malicious shards with AUROC 1.0 across tested setups, but stealthy low-magnitude attacks slip under real hardware noise.
Peer-to-peer Large Language Model (LLM) inference slices a model across machines owned by different people. It is cheap and decentralised, but there is a catch: any shard in the chain can nudge the activations it forwards and poison the final answer. Recomputing on trusted hardware spots that, yet the cost is brutal and exact-match checks fall over because benign hardware noise already shifts numbers around.
How the canaries work
This paper takes a simple, sneaky idea and makes it measurable. Precompute a reference activation for a small, secret set of canary prompts, then quietly mix those canaries into normal requests. At run time, shards execute in lower precision (fp16), return their activations, and the verifier measures how far a canary’s live activation drifts from the stored reference using relative L2 distance. You do not fix a threshold. You rank shards by how often they look “too far” on canaries and score the separation with AUROC.
On the test bench, that ranking is razor sharp: the detector reaches AUROC 1.0 on all 72 factorial configurations where it is active, cleanly putting the malicious shard above every benign shard for every canary. It holds across GPT-2 and Pythia-style blocks, two datasets, pipeline depths from 2 to 12 blocks (with a 24-block subset), different attacker positions, and even up to four colluding shards. Choice of distance matters: cosine distance misses pure rescale attacks, while relative L2 still flags them. You do not strictly need a pristine fp32 reference either; a noisy one still works.
Attacker’s angle
The Achilles’ heel is the noise floor. When benign hardware variation climbs toward the attacker’s perturbation, the AUROC slides toward 0.5. With the modelled floor around 0.002, detection holds down to a similar tamper magnitude and breaks below it. That sets a hard trade-off: stay under the noise and you cause essentially no harm; push hard enough to matter and you light up the detector.
Evasion is feasible in principle but constrained in the study. If an adversary can profile the pool to estimate the benign floor, they can pick a tiny perturbation that hides. If they can fingerprint canaries, they can spare them. The evaluation assumes neither is allowed: attackers tamper every query at a fixed strength, and canaries stay secret. Real hardware noise is also more structured than the Gaussian model here, and the models tested cap out at 410M parameters. Whether these clean separations survive at larger scales and in adaptive, messy pools is the live question. The method is practical and cheap to run; the politics will be in keeping canaries secret and measuring the true noise floor in the wild.
Additional analysis of the original ArXiv paper
📋 Original Paper Title and Abstract
Integrity of peer-to-peer distributed LLM inference under malicious nodes
🔍 ShortSpan Analysis of the Paper
Problem
This paper studies output integrity for peer-to-peer distributed LLM inference, where model layers are split across many independently owned nodes. Any node can maliciously alter the intermediate activations it forwards, corrupting final answers. Exact-match or recompute-based defences detect tampering but are costly and brittle because benign nondeterminism from heterogeneous hardware, floating-point rounding and low-precision execution already produces variation in activations. The paper asks whether a cheaper, practical detector can reliably single out a malicious shard despite ordinary benign drift.
Approach
The authors transplant known-answer canaries to intermediate activations. A verifier precomputes fp32 reference activations for a small set of secret canary inputs and mixes them into live traffic; canaries are drawn from the same input distribution so they are indistinguishable at serve time. Live passes run in fp16 and each shard returns activations. For each canary and stage the verifier computes a relative-L2 mismatch between live activation and the stored reference. Detection is cast as a threshold-free ranking problem scored by AUROC: a malicious shard should consistently outrank benign shards on mismatch. The evaluation uses transformer blocks from GPT-2 and Pythia variants, two datasets (LAMBADA and WikiText-2), six tampering strategies (Gaussian perturbations at several magnitudes, sign flip, rescale), 100 canaries per configuration, a factorial of 216 configurations and five substudies for 408 total configurations. Two ablations test the role of a stored reference and use of cosine distance versus relative-L2.
Key Findings
- The full detector attains AUROC 1.0 on all 72 factorial configurations where it was active, correctly ranking the malicious shard above every benign shard for every canary; the redundant-recompute ceiling also reaches 1.0.
- The signal is robust across architectures, datasets, pipeline depth (2 to 12 blocks, including a 24-block model subset), attacker position and up to four colluding malicious shards; these axes do not reduce AUROC when the benign noise floor is sufficiently below the tamper magnitude.
- Ablations show the stored fp32 reference is not essential (a noisy reference still works) but the choice of distance matters: cosine distance fails to detect pure rescale attacks while relative-L2 still detects them.
- Detection breaks down when benign noise approaches the tamper magnitude. With the modelled noise floor σhw=0.002 the detector holds down to tamper magnitude ε≈0.002 and fails below that; when σhw rises toward ε the AUROC degrades toward 0.5.
- There is a practical evasion trade-off: attacks with ε below the noise floor evade detection but cause essentially zero harm to model accuracy, whereas harmful attacks require ε orders of magnitude above the floor and are detected.
Limitations
The experimental noise model uses isotropic Gaussian perturbations while real hardware nondeterminism is structured and correlated; results therefore depend on where the true benign floor sits relative to tamper magnitudes. Experiments use models from 124M to 410M parameters; behaviour on very large deployed models may differ. The adversary model is constrained: attackers commit to a tampering magnitude and tamper on every query and are not allowed to probe the noise floor or to fingerprint and selectively avoid canaries. Secrecy of which requests are canaries is assumed. These factors limit immediate generalisability to operational pools without further profiling and counter-adaptive analysis.
Implications
Offensive implications are twofold. First, an attacker can evade this defence by profiling the pool to estimate the benign noise floor and then selecting tamper magnitudes below that floor, or by developing methods to distinguish and spare canaries if they can fingerprint them, enabling persistent undetected corruption at low amplitude. Second, attackers facing detection must choose between making large, detectable perturbations that degrade outputs or making tiny, stealthy perturbations that do not affect utility. Scaling to larger models and adaptive adversaries could change that trade-off, so attackers may prioritise reconnaissance and selective tampering.