FR
live

One misconfigured default route took 27% of Solana’s stake offline for 33 minutes

On August 12, 2026, a malformed default route propagated by a route reflector at hosting provider TeraSwitch knocked 28.83% of staked SOL offline, within 4.51 points of Solana’s finality halt threshold. The lesson goes beyond blockchains: measure your ASN concentration and actually test your failover.

A fiber-optic patch cable pulled partway out of a dense patch panel in a dark server room, its exposed connector tip catching a faint amber glow.

August 12, 2026. 28.83% of staked SOL dropped offline within minutes. 04:16:15 UTC, service restored. For roughly half an hour, Solana came within 86% of the threshold where no transaction gets finalized — not because of a bug in its validator client, but because of a misconfigured default BGP route at its hosting provider, TeraSwitch.

The incident never caused a halt, which is exactly what makes it worth an SRE or CISO’s attention: a single malformed route, pushed through a route reflector, took down 12 data centers across two continents and cut 94% of the stake hosted on one autonomous system. The fragility was not in the blockchain — it was in the BGP underneath it.

What happened: a default route stripped of its attributes

The failure, reconstructed by liquid staking protocol Marinade Finance and later confirmed in a TeraSwitch postmortem, boils down to one propagation chain. A default route originated at the provider’s Miami (MIA1) site travelled across its backbone with its metric and BGP communities stripped, and an AS-path containing only TeraSwitch’s own ASN — AS20326.

A route reflector in Amsterdam (AMS2) then pushed that altered route into the provider’s European and Asian markets. Downstream edge routers read it as locally originated, preferred it over their own valid local default, and advertised it into the data-center core — which rejected it as invalid. With no acceptable default installed, the affected site fabrics stopped forwarding traffic to their own edge routers.

Twelve sites lost reachability — LON1, AMS1, AMS2, AMS3, DUB1, DUB2, FRA2, SGP1, SGP2, TYO1, TYO2, TYO3 — while North American sites stayed up. Engineers spotted the malformed route within about ten minutes and pulled MIA1 from the backbone; full traffic restoration was logged at 04:16:15 UTC, closing a window of roughly 30 to 33 minutes.

28.83% delinquent stake, 4.51 points from a halt

Solana consensus counts stake, not servers. During the window, 28.83% of staked SOL went delinquent, against a 33.34% threshold above which the network stops finalizing blocks. In other words, Solana got 86% of the way to losing finality, with 4.51 percentage points of headroom — about 19.9 million SOL.

By headcount the event looked smaller: 597 of 699 validators kept voting, so 102 dropped (about 15%). But the stake figure — 28.83% — is the one consensus reads. Blocks kept being produced and transactions kept landing, with no halt and no rollback. Solana’s status page logged no incident, consistent with a chain that degraded but never stopped.

The financial cost stayed minimal: 333 SOL in missed staking rewards, about $25,600, absorbed by validator bonds at epoch end. SOL traded around $76.31 to $76.46 during and after the incident.

Nobody failed over: the real problem is operational

The most uncomfortable finding is operational, not architectural. Marinade measured how 74 validators behaved during the blackout: only 3 failed over cleanly to standby infrastructure. 59 validators holding 80.2 million SOL came back within the same narrow window across Amsterdam, Frankfurt and Tokyo — the signature of passively waiting for BGP reconvergence rather than actively migrating.

Even Helius, one of the ecosystem’s larger infrastructure operators, stayed offline for the full 33 minutes after its backup systems failed to activate. Marinade’s framing is blunt: “Nobody failed over. They sat there until the routing reconverged.”

That is lesson number one for any operator: a failover you never test does not exist. The 4.51-point margin that saved finality was set by a third party’s response time — TeraSwitch — not by the resilience of the operators holding the stake.

27.34% of stake on one ASN: concentration as a structural risk

The number that turns this from an outage into a structural warning is 27.34%: AS20326, TeraSwitch’s autonomous system, carries 118,890,767 SOL — more than a quarter of all staked SOL. That already exceeds the 25% per-autonomous-system cap in the Solana Foundation Delegation Program, and 94% of that stake went dark within the same few minutes.

The instructive part: the cap did its job on the stake it controls. Marinade checked the 82 validators on AS20326 and found zero foundation delegation — “not one SOL of 24.5M”. The concentration was assembled by market choice: price, latency and operational convenience pulled independent validators onto the same fabric.

Marinade then turned the same measurement on itself: two-thirds of the stake its allocation model distributes sits on four autonomous systems, with AS395201 at 36.94% — higher than the share that nearly stalled finality. Its own verdict: “Nobody should be comfortable with that, us included.”

What this teaches a network that has nothing to do with crypto

It would be easy to file this under crypto. It is the opposite: it is a BGP problem that happened to find a very measurable victim.

First, RPKI and Route Origin Validation (ROV) would not have stopped it, as we explained in our RPKI and ROV guide. Those mechanisms catch origin hijacks — a prefix announced by an ASN that does not legitimately own it. The faulty route carried TeraSwitch’s own ASN: it was not a hijack but an internal corruption propagated by a reflector. The fix is configuration hygiene: do not re-originate defaults, do not strip attributes, and control what your reflectors propagate.

Second, ASN concentration is a measurable and neglected risk. A single ASN carrying 27% of your consensus weight — or of your production workloads — turns an isolated configuration error into a systemic incident. The metric to watch is not “how many providers” but “what fraction of capacity goes down if one ASN fails”.

Third, passive failover does not count. Redundancy that “comes back on its own once routing reconverges” is not redundancy: it is a disguised dependency on a third party’s remediation speed.

Put together, the three lessons form a single discipline: treat the network underneath your workload as part of your availability model, not as a given. The fact that Solana’s own status page logged no incident — because there was no halt to log — is the trap: an infrastructure near-miss that never trips an alert is exactly what your monitoring will not catch. Ask your providers for their route-propagation controls, measure your ASN concentration, and fail over on purpose, not on faith.

Verdict

If your availability rests on one host or one ASN, measure the fraction of capacity it concentrates and ask it — as TeraSwitch did — for a public postmortem on its route-propagation controls.

If you run critical services, test failover for real: not a paper scenario, but an actual cutover proving your standby activates without waiting on BGP reconvergence. The August 12 incident shows a 33-minute window without finality is a plausible scenario, and that the margin separating you from an outage is often set by someone else.

For Solana, the network survived a live stress test with no halt and no loss of funds — but the structure that put it 4.51 points from the cliff is still in place: 27% of stake on one ASN, and validators that wait for reconvergence instead of failing over.

References

The cyber brief, every Tuesday

The flaws that matter and the patches to apply, in a ten-minute read.

No spam. One-click unsubscribe.
read next

On the same topic

HPE closes its $14 billion Juniper acquisition after two and a half years

On August 13, 2026 a federal judge approved the settlement between HPE and the US Department of Justice, ending a two-and-a-half-year regulatory saga over Juniper Networks. For network teams, the HPE–Aruba–Juniper combination redraws the enterprise switching and Wi-Fi market against Cisco and Arista.

← Back to the feed

Type at least two characters.

navigate open esc dismiss