Did you know that you can navigate the posts by swiping left and right?
Multihop entanglement : what a quantum network does when a fiber breaks
27 Apr 2026
. category:
networking
.
Comments
#networking
#fiber
#quantum
every transport network has an answer to the backhoe. sonet had automatic protection switching and a 50 ms budget for it. sdh had the same. optical mesh networks have pre-computed restoration paths, shared spare capacity and a whole literature on how to place it. the assumption underneath all of it is that fiber gets cut, and the network is supposed to notice and keep going.
quantum networks have had no answer at all. every entanglement distribution experiment before this one ran on exactly one lightpath between the source and each user. cut it and the service stops until somebody walks to the patch panel.
this paper is the first one to fix that. Alshowkan, Lukens, Lu and Peters, “Resilient Entanglement Distribution in a Multihop Quantum Network,” Journal of Lightwave Technology 43(19), 9016 (2025) (arXiv). six users, three buildings on the Oak Ridge campus, one entangled photon source, and a controller that watches for a dead link and reroutes the spectrum through a different building.
it works. it also takes five seconds, costs about 10 dB per hop, and in one of the two recovery scenarios the restored links come back at a fidelity of 0.76. that last number is the one i want to spend time on, because the paper reports it accurately and then moves on, and i think it is the most important thing in the paper.
if you’ve read the flex grid post, this is the same group’s network four years later, with a control plane bolted on. if you haven’t, the short version is below.
what came before
one box makes entangled photon pairs. every user has a fiber to that box. the pairs come out over a wide band of wavelengths at once, and because the two photons of a pair always add up to the pump energy, wavelength is the address:
ω_signal + ω_idler = ω_pump
hand two users energy-matched slices of that spectrum and they share entanglement. hand them mismatched slices and they share nothing. so provisioning who-talks-to-whom is a spectrum allocation problem, and a wavelength selective switch, the same liquid-crystal-on-silicon part that lives in every ROADM, is the thing that does the allocating.
that was the 2021 result. this paper takes that architecture and asks what happens when the fiber underneath it breaks.
the network
three buildings on the ORNL campus, one subnetwork each.
| Building | Users | Detectors | Fiber to A |
|---|---|---|---|
| A | Alice (A1), Avery (A2) | SNSPD, >81% | the source is here |
| B | Bob (B1), Bailey (B2) | free-running APD, 20% | 250 m |
| C | Charlie (C1), Casey (C2) | SNSPD, >81% | 1.2 km |
plus a 930 m fiber between B and C. three buildings, three fibers, a full mesh, which at N=3 is also the minimum for any protection at all.
the source in building A is a continuous-wave laser at 779.4 nm putting about 25 mW into a type-II PPLN ridge waveguide, giving polarization-entangled pairs in the state |Ψ+⟩ ∝ |HV⟩ + |VH⟩ across roughly 310 GHz. WSS1 carves that into eight pairs of frequency-correlated 25 GHz channels centered on 192.3125 THz, which puts them on the standard ITU grid, channels 21.25 to 25.00.
two of WSS1’s outputs go to the local users. two more go to the patch panel and out to B and C, where WSS2 and WSS3 act as local hubs and split the arriving spectrum among their own users.
three things make it a network rather than an experiment:
a clock. a rubidium frequency standard disciplines a White Rabbit switch, which feeds a White Rabbit node in each building over 1310/1490 nm SFPs. those nodes hand out 10 MHz (doubled to 20) and a pulse per second to the time-to-digital converters that timestamp every detection. the coincidence window is about 1 ns. if you want the argument for why the clock is the hard part, it’s in the coexistence post.
a control plane. an SDN controller programs the WSSs and a set of MEMS switches. the quantum data plane is all-optical: no optical-to-electrical conversion anywhere, so nothing in the path measures or copies the photons.
redundant fiber and MEMS switches. building B has a 2×1 MEMS switch and a circulator; building C has a 2×2. these are what let the controller pick which physical fiber feeds the local WSS, and they are the entire protection mechanism.
one thing worth flagging early: the classical and quantum signals run on separate fibers. that sidesteps Raman scattering and filter crosstalk entirely, and the paper says so. it is also a luxury that a real deployment on leased fiber will not have.
what a hop costs
in normal operation, the total loss from the input of WSS1 to a user’s detector is 7.25 dB (A1) and 8.51 dB (A2) for the local users, and 16.6 to 17.6 dB for everyone in B and C.
when a link fails and traffic is rerouted through the other building, C1 goes from 17.2 dB to 27.5 dB. B2 goes from 16.6 dB to 28.0 dB. the paper’s summary: about 10 dB per extra hop.
almost none of that is fiber. 250 m of SMF-28 is 0.05 dB. the 10 dB is the insertion loss of a second WSS, plus the MEMS switch, plus connectors. on a campus this size the network is entirely component-limited, and the fiber is free.
that matters because of how a coincidence rate is built. a link needs both photons to arrive. so the rate scales as the product of the two arms’ transmission, and whether one photon or two take the detour changes the answer by an order of magnitude.
the C1–C2 link is the one to look at. both Charlie and Casey sit in building C, so when C’s feed is rerouted through B, both photons of every pair take the extra hop. their singles rates drop by 10.9x and 11.2x, and the coincidence rate goes from 48.1 to 0.33 per second, a factor of 146 where the product of the two singles reductions predicts about 123.
for comparison, A1–C1 has only one photon on the detour and loses a factor of 10.
that asymmetry is a routing fact, not a physics curiosity. an intra-building link pays every hop twice. a controller choosing a restoration path should know that the users whose service degrades worst are the ones on the far side, talking to each other.
the baseline
before breaking anything, they characterize the star topology: WSS1 in the middle, direct fibers to B and C. six links, tomography on all of them at once, 36 polarization projections each, Bayesian state estimation.
| Link | Type | Fidelity to |Ψ+〉 | Coincidences/s |
|---|---|---|---|
| A1–A2 | intra, building A | 0.9504(8) | 866.5(7) |
| B1–B2 | intra, building B | 0.887(8) | 9.16(7) |
| C1–C2 | intra, building C | 0.932(3) | 48.1(1) |
| A1–B1 | inter | 0.965(8) | 43.4(4) |
| A2–C1 | inter | 0.978(2) | 246.6(1) |
| B1–C1 | inter | 0.884(2) | 9.8(2) |
nothing surprising. the links with an APD user in building B are the slow ones, which is a detector problem, not a network problem. the good links are up around 0.97.
note the rates, though. the fastest link on this network moves 866 pairs per second. the slowest moves 9. an ethernet frame is 12,000 bits. these are not the same units as anything a network engineer is used to holding in their head, and it is worth recalibrating before reading the recovery numbers.
breaking it
the failure model is a single building-to-building fiber going away. with three buildings in a triangle, any one of the three can be lost and the other two still reach each other the long way round.
detection is a threshold. the SDN application layer sets a minimum singles count for each detector, chosen from the dark count and background level. when a detector’s singles fall below it, the controller assumes the upstream fiber is gone and starts the reroute: WSS1 is told to send building C’s slots to building B instead, WSS2 is told to pass the extra spectrum through to its output fiber, and the MEMS switches flip to put the right fiber on the right WSS input.
losing A→C
the controller sends building C’s spectrum to building B, and building B passes it on. C’s users are now two hops from the source.
fidelities come back at 0.947, 0.897 and 0.909, within a few points of their direct-path equivalents. rates are the casualty: A1–C1 lands at 23.9/s against 246.6/s for the comparable direct link, B2–C2 at 2.94/s against 9.8/s, and C1–C2 at 0.33/s against 48.1/s.
this is the good case, and it is worth being clear about why. the loss went up 10 dB and the fidelity barely moved, because loss is the one impairment that does not touch the quantum state. you lose photons, you don’t corrupt the ones that survive. attenuation costs you rate and nothing else.
that is a genuinely nice property, and it is the reason a transparent optical reroute is a reasonable thing to do at all. it also runs out.
losing A→B
the other direction is where it gets interesting.
building B’s users are the ones with the 20% APDs. rerouting them through C would put an already-slow link 10 dB further down, and the paper decided that was not survivable at the normal allocation. so instead of the balanced eight-channel plan, they used what they call priority allocations: give the entire band to one link at a time.
two configurations, mutually exclusive. all eight channels to A1–B2, then all eight to B2–C2. rates come back at 17.5/s and 4.92/s, which is within a factor of two or three of the direct links. the loss got compensated.
fidelity is 0.765(8) and 0.76(2).
the trade they made
the paper explains the fidelity drop in one sentence: the larger bandwidth increases multipair emission noise inside the coincidence window.
that is worth unpacking, because it is the most useful thing in the paper and it is easy to read past.
a continuous-wave pumped SPDC source does not emit one pair at a time on request. it emits pairs at random, and occasionally two pairs land inside the same 1 ns coincidence window. when that happens, the photon Alice detects and the photon Bob detects came from different pairs, and they are not entangled with each other. they are an accidental coincidence dressed up as a real one.
the scaling is unforgiving. allocate a link n channels instead of one and the true coincidence rate goes up like n, because you are collecting more of the same pairs. but the accidentals go up like n², because an accidental needs one photon from each of two independent singles streams, and both singles rates scaled with n:
true coincidences ∝ n
accidental coincidences ∝ n²
so the ratio of garbage ∝ n
so bandwidth buys rate linearly and buys noise quadratically. the paper traded fidelity for rate, and the exchange rate is set by the pump power and the coincidence window, not by anything in the network.
they were right to make the trade in that situation. the alternative was leaving building B dark. but it was made by hand, as a configuration choice, and nothing in the control loop knows the exchange rate exists.
where 0.76 sits
the paper reports log-negativity of 0.62(2) and 0.61(3) ebits for the two recovered links and notes, correctly, that the states are still entangled and still useful.
model the noise as depolarizing, which for this data is a good enough approximation, and a state with fidelity F to a Bell state has
p = (4F − 1) / 3 depolarizing parameter
E_N = log₂( (3p + 1) / 2 ) log-negativity, in ebits
at F = 0.765 that gives E_N = 0.614. the paper measured 0.62(2). the model is calibrated.
so run it the other way and ask where the thresholds are. the same model says a Bell inequality is violated only when p > 1/√2, which is
F > 0.780
both recovered links land just under it. 0.765 and 0.76.
that is not a catastrophe. log-negativity is positive, the entanglement is distillable, and plenty of protocols work fine below the CHSH line. but device-independent protocols do not, and any user who was relying on a Bell violation as their security argument just silently lost it. the network came back up. the service did not come back the same.
and the network cannot tell. the health signal in this system is a singles count threshold. singles were fine. they were better than fine, because the recovery deliberately widened the band to push them up.
look at the shape of that second panel. raw coincidence rate is a straight line and ebits per second is not. the paper’s own metric of choice from earlier work was ebits/s, and on that metric the last few channels are nearly free of benefit while costing all of the fidelity margin.
five seconds
the recovery time, end to end, including detection and the sequential switching of the WSS and MEMS devices, is approximately 5 seconds.
for context, the number a transport network has been engineered around since the 1980s is 50 ms. that figure comes from how long a voice circuit can be interrupted before it drops, and it survived into SDH, into optical mesh restoration, and into every carrier SLA written since. five seconds is two orders of magnitude past it.
the paper doesn’t break the five seconds down, and that’s the single most useful missing number in it. the components have very different speeds:
- MEMS switches are sub-millisecond to a few milliseconds. not the problem.
- LCoS wavelength selective switches take somewhere between a hundred milliseconds and a few seconds to load a new profile, because the liquid crystal has to physically settle. very likely a large share.
- detection needs enough integration time on a singles counter to distinguish a real failure from a statistical dip. on the APD users in building B, sitting at a few thousand counts per second, a confident threshold crossing is not instant.
- polarization re-settling after the photons take a physically different fiber. the paper does not say whether the analyzers needed re-optimizing, and if they did, that could dominate everything else.
each of those has a different fix, and you cannot pick one without knowing which term dominates. an oscilloscope, a counter and a long afternoon would settle it.
what’s missing
the paper is a first, and firsts get graded generously. these are the things i’d want before calling this a protection scheme rather than a demonstration of one.
the failure detector measures the wrong thing. a singles-count threshold is a loss-of-light alarm. it catches a cut fiber. it does not catch a link that is passing photons and no longer passing entanglement, which is exactly what a polarization drift, a source degradation, or a widened allocation produces. the A→B recovery is the proof: the network restored a link to 0.76 fidelity and had no mechanism that could have noticed. the health metric and the service metric are different quantities, and only one of them is being watched.
and it cannot localize. a threshold crossing at C1 says something upstream is broken. it does not say what. classical transport spent thirty years building an alarm hierarchy (loss of signal, loss of frame, alarm indication signal, remote defect indication) precisely so that one fault raises one alarm at the right place instead of every downstream node reporting a failure simultaneously. a quantum network with more than three buildings will need the same discipline, and there is nothing here yet.
the protection is 1:1 and dedicated. every building pair has its own fiber, standing by. that is fine for a triangle. it is N(N−1)/2 fibers for a full mesh, and the entire history of optical network design is the story of moving away from dedicated spare capacity toward shared pools. the quantum version has a shared resource classical networks don’t: spectrum. the spare capacity to pool is unallocated channels on the switch, not dark fiber. nobody has written that problem down.
shared risk is unaccounted for. the paper notes the fibers “pass through several telecommunication rooms” before reaching their destination. if the A→B and A→C fibers leave building A through the same duct, they are one fiber wearing two names, and the protection is decorative. classical planning calls this a shared risk link group and treats an SRLG-diverse path as the minimum bar for calling something protected. establishing it needs a duct map and, realistically, OTDR traces to confirm the physical routes actually diverge.
the source is a single point of failure. one PPLN waveguide in building A feeds the entire network. every fiber is protected and the thing generating the entanglement is not. in classical terms they hardened the links and left the head end alone.
degraded-mode policy is manual. in the A→B recovery, the two priority allocations are mutually exclusive: the network can serve A1–B2 or B2–C2, not both, and which one is a human decision made at configuration time. a real network needs this as policy, with committed rates, preemptable classes and an admission decision that comes with a reason attached, rather than as a pair of profiles someone loads by hand.
one fault, and only one. the triangle survives any single cut and no double cut. that is a fine place to start, and it is worth stating explicitly rather than leaving “resilient” to do the work.
“multihop” is doing double duty. in most of the quantum networking literature, a hop implies a repeater: a node with a memory that does an entanglement swap, and whose whole purpose is that the loss budget starts over on the other side. nothing here swaps. the photon physically transits building B’s WSS on its way to C, and pays 10 dB for the privilege. this is an optical bypass through a ROADM, which is a completely legitimate and useful thing, but it is a different thing. the paper’s discussion section is careful about this. it says the term refers to architectural capability rather than a fixed number of hops. the title is not so careful.
what i’d build next
1. publish the five-second breakdown, then attack the biggest term. detection, WSS settling, MEMS, polarization re-lock. four numbers.
2. pre-load the restoration profile so the slow device is out of the failure path. if the WSS is the bottleneck, don’t reprogram it during an outage. provision the restoration allocation on a spare WSS output at commissioning time and leave it loaded, so that recovery is a MEMS flip and nothing else. you pay in spectrum, because the standby allocation occupies channels that carry nothing until they’re needed. that is precisely the 1+1 versus 1:1 trade from classical protection, translated into the spectral domain, and as far as i can tell nobody has posed it that way. the question “how much spectrum should a network hold in reserve for restoration” has a clean formulation and no answer in the literature.
3. alarm on entanglement, not on photons. full tomography is 36 settings and hours, so it cannot be the health signal. but you don’t need it. a two-basis correlation check, or a coincidence-to-accidental ratio monitor, is cheap and continuous and would have caught the 0.76. better still, the in situ process tomography work gives a live quantum map of a channel with a sliding window, including noise the classical polarization tracker cannot see. that is a control-plane input, and it is sitting unused. the ORNL group has also started on what a standardized quantum network metrics set should even contain, which is the other half of the same problem.
4. put the multipair term in the objective function. the allocator maximizes rate under a fidelity constraint. but as shown above, the fidelity of a link is a function of how much bandwidth you gave it, and that coupling isn’t in any of the routing-and-spectrum-allocation formulations. it is one analytic term. adding it turns “restore the link” into “restore the link to a stated service class”, which is the difference between the two.
5. formulate shared restoration in spectrum. classical spare capacity allocation is a well-studied problem with decades of heuristics. the quantum version has a different shared resource and an extra constraint, and it is a straightforward port for somebody who knows both literatures.
6. protect the source. two sources, in different buildings, and a controller that treats source diversity as another dimension of the allocation. the 2026 routing and spectrum allocation work already handles multiple sources and non-colliding frequency bin assignment, so the solver side exists. what’s missing is the failover.
7. put the coexistence noise back in. this network keeps quantum and classical on separate fibers, which is what makes the reroute clean. share the fiber and the noise floor on a restoration path depends on which classical channels are lit on the glass you just rerouted onto, because spontaneous Raman scattering dumps noise across tens of THz. the restoration decision then depends on the classical traffic matrix, which is a coupling nobody has written down. the same group’s coexistence work has the noise model.
8. and eventually, hops that actually reset the budget. the 10 dB per hop is a consequence of transparency. a node with a quantum memory that performs an entanglement swap breaks the multiplication instead of compounding it, and the routing problem changes character completely: you would route toward nodes that have memories rather than away from extra hops. that used to be a lab-only capability. in february 2026 Qunnect and Cisco ran entanglement swapping over 17.6 km of deployed metro fiber between Brooklyn and Manhattan at 5,400 swapped pairs per hour with polarization fidelity above 99%, using independent sources rather than a shared master laser. 5,400 an hour is 1.5 a second, against 9.8 to 866 a second for the direct links here, so a swap is still two to three orders of magnitude slower than just sending the photon through. transparent bypass is the right answer today. it will not be forever, and the control plane built here is the part that carries over.
the short version
- this is the first quantum network that reroutes entanglement around a broken fiber instead of waiting for someone with a screwdriver. six users, three buildings, one source, an SDN controller driving wavelength selective switches and MEMS switches.
- a rerouted hop costs about 10 dB, almost all of it component insertion loss rather than fiber. at campus scale the glass is free and the switches are not.
- because a coincidence needs both photons, one photon on the detour costs about 10x and two photons cost well over 100x. the C1–C2 link fell from 48.1 to 0.33 coincidences per second, a factor of 146.
- loss costs rate and not fidelity, which is why a transparent optical bypass works at all. attenuation removes photons; it does not corrupt the survivors.
- in the harder recovery they gave one link the whole band to compensate for the loss. rate came back within a factor of two to three. fidelity fell to 0.76, because bandwidth buys true pairs linearly and accidental pairs quadratically.
- 0.76 is just below the ~0.780 needed for a Bell violation under a depolarizing model. still entangled, still distillable at ~0.62 ebits, but a device-independent user silently lost their security argument and the network had no way to know.
- recovery takes about 5 seconds, against the 50 ms that transport networks have been engineered around for forty years. the paper doesn’t say which component owns the five seconds.
- the failure detector is a singles-count threshold. it catches a cut fiber and cannot catch a link that passes photons but no longer passes entanglement, which is exactly what the recovery produced.
- protection is dedicated 1:1 on pre-built fiber, single-fault only, with no shared risk analysis, no source redundancy, and degraded-mode priorities chosen by hand.
- “multihop” here means transparent optical bypass, not repeater hops. nothing swaps, so nothing resets the loss budget. the paper’s discussion says this; the title doesn’t.
the thing i keep coming back to is that this paper’s real contribution isn’t the reroute. it’s that it forces the question of what “up” means on a quantum link. on a classical link, up is unambiguous: light is arriving, frames are checksumming, done. here the network restored service, reported healthy, and delivered a state that a whole class of protocols can no longer use. carrying photons and carrying entanglement are two different conditions, and this is the first paper that puts a measured number on both of them at the same time. everything an operations team would want next, the alarms, the thresholds, the service classes, the escalation, follows from taking that distinction seriously.