Did you know that you can navigate the posts by swiping left and right?
holding entanglement together on a wire in the wind
20 Jul 2026
. category:
networking
.
Comments
#networking
#fiber
#quantum
i have now cited this paper twice in passing, once in the fiber types post as the reason polarization drift is a solved problem, and once in the eight-users post under “what happened since.” it is out properly now, so it is time to write it up.
disclosure before anything else: i am on it. Shi et al., “Entanglement distribution over a polarization-stabilized aerial fiber,” J. Opt. Commun. Netw. 18, 813 (2026), also on arXiv. that changes what this post can be. i am not going to pretend to review it from outside, and i am going to be more specific than a reviewer could be about which numbers were decisions and which were things the fiber did to us. what it does not change is the format: there is still a “what i’d push on next” section, and some of it is aimed at my own paper.
the one-line version is that we sent one half of an entangled photon pair 62 km from NIST in Gaithersburg to the University of Maryland in College Park, over a fiber that is about 70% strung between poles, and it worked for a day. Oliver Slattery, one of the co-authors, described the link to the NIST press office as “about as bad a connection as you can possibly have,” which is roughly the spirit in which it was chosen.
this post is the long version. the first section is background for people who do not do quantum. if you know what a Bell violation is, skip to what the wind does.
the quantum part, for people who don’t do quantum
polarization entanglement, in one paragraph
a source makes two photons at once. the pair is prepared in the state
|Φ+⟩ = (|HH⟩ + |VV⟩) / √2
which says: if you measure both photons in the horizontal/vertical basis you will always find them the same, and, the part that matters, that also holds in every other basis. measure both at 45°, still the same. this second half is not something a pair of classically correlated photons can do, and the whole apparatus of quantum key distribution and entanglement swapping rests on it.
we send one photon of each pair down the fiber. the other stays home.
why polarization, and why it is the fragile choice
there are other places to hide a qubit in a photon. time-bin encoding survives fiber beautifully, which is why long-distance records tend to use it. frequency-bin is having a moment. polarization is the one that is trivial to make, trivial to measure with a waveplate and a beamsplitter, and trivial to feed into any of the standard protocols. it is also the one the fiber attacks directly.
a perfect fiber would be a perfect cylinder of perfectly isotropic glass and would leave polarization alone. real fiber has an elliptical core, sits in a duct under load, gets bent around a pole, and heats unevenly. all of that is birefringence, which means the two polarization components travel at slightly different speeds and pick up a relative phase. i covered the mechanism in the fiber types post; the short version is that a length of deployed fiber applies some arbitrary rotation to the polarization state of everything passing through it.
if that rotation were fixed, this would be a non-problem. you would measure it once, put a compensating rotation at the far end, and never think about it again. it is not fixed. it moves with temperature, with stress, with anything that touches the cable.
buried fiber is a different animal
this is the part worth internalizing if you come from the classical side, where the aerial-versus-buried decision is about cost, permits and backhoes.
buried fiber sits in a duct at a stable temperature with nothing touching it. its polarization transformation drifts on a scale of hours, sometimes longer. Wengerowsky and co-authors distributed polarization entanglement over 192 km of deployed fiber with no active stabilization at all, which tells you how gentle a buried link can be.
aerial fiber is hanging in the air. the sun heats one side of the span. wind moves it. a truck goes past and shakes the pole. birds land on it, which is a sentence that appears in the NIST press release and is not a joke. the same transformation that took hours to move underground now moves in seconds.
so the interesting question is not “can you send entanglement over fiber,” which was settled years ago. it is whether the fraction of the installed base that is strung between poles is usable at all, because in most metro areas that fraction is not small and you do not get to choose.
what the wind does
the link is 62 km of deployed fiber between NIST Gaithersburg and UMD College Park, roughly 70% aerial. end to end loss is about 18 dB. of that, roughly 12.4 dB is the fiber doing what fiber does at 0.2 dB/km, and the other 5.6 dB is splices, patch panels and switches, which is a very ordinary ratio for a real link and worth remembering the next time somebody quotes you a length-times-attenuation budget.
before touching any of the quantum hardware we did the boring characterization: inject polarized probe light, watch the Stokes parameters come out the far end, and see how long the state stays put.
the answer depends on what time it is. at night the output polarization is stable for tens of minutes. during the day it walks tens of degrees away in a matter of seconds, and the fidelity of the transmitted state falls below 95% in under 20 seconds.
the factor between day and night falls straight out of those two published numbers, and it is about 60.
the daytime average is not the number that matters, though. at midday the state takes about 11 seconds to cross the 98% threshold that triggers the compensator, so the 3 second window we leave open for entanglement has a comfortable 3.5x of margin. under a gust it crosses in about 2 seconds, which is inside the window. the link is not marginal on a typical afternoon. it is marginal specifically during the events that are also making the compensator time out, and that pattern comes back twice more in this post.
a pilot tone and an inverse
the fix is the one you would guess. send a classical reference signal through the same fiber, measure what the fiber did to it, and drive an inline polarization controller until the damage is undone.
the specific arrangement here is a pair of automated polarization compensation modules built by Qunnect: an injector at the NIST end and a compensator at the UMD end.
- the injector launches about 0.5 mW of laser light into the fiber, cycling its polarization through a sequence of six designated reference states.
- the compensator measures the six states as they arrive, on a polarimeter, and computes the fidelity of each against what was sent.
- the minimum of those six fidelities becomes the error signal. a gradient descent drives the polarization controller to push it back toward one.
six reference states is more than the job strictly requires. two non-orthogonal states are mathematically sufficient to pin down the transformation. more states means a better-conditioned estimate and a slower modulation cycle, and six is where that trade bottomed out for this hardware. this is the kind of thing that reads as arbitrary in a paper and was in fact a couple of afternoons.
the loop is a threshold machine, not a continuous servo:
- check the fidelity. above 98%, do nothing, release the fiber immediately. that check takes about 50 ms.
- below 98%, run the descent until fidelity exceeds 99%, or until 55 seconds have gone by, whichever happens first.
- either way, hand the fiber back for a 3 second window, then check again.
the 55 second timeout is there so that a session cannot swallow the link forever during a bad stretch. it also, as we will get to, ends up being the most interesting number in the experiment.
why the pilot has to sit on top of the photons
this is the constraint that shapes the whole design, and the one i have the most to say about.
the pilot tone is only useful if the fiber does the same thing to it that the fiber does to the single photons. birefringence is wavelength dependent, so that stops being true as you move them apart. we measured it: six polarization basis states across the 62 km at wavelengths from 1545 to 1555 nm, extract the transformation matrix at each, and compare.
the fidelity between transformations falls off quickly with detuning. so the pilot is generated at 1549.32 nm, the same ITU channel as the idler photons it is protecting.
that is not quite what the paper says, and the difference is the part i want to argue with.
the paper’s stated rule is that the reference should sit within 1 nm of the photons. that is sound advice. but a nanometer of detuning costs about 0.005 in transformation fidelity, and when you push that through to a measured S-value it is worth roughly 0.01. you have to get out to about 4 nm before the leftover rotation alone is enough to drag S through the classical bound. one ITU channel away, 0.8 nm, is essentially free.
so the wavelength argument does not on its own force the pilot into the same channel as the photons. and that matters, because being in the same channel is exactly what forces the time multiplexing, and the time multiplexing is what costs 7.2% of the link.
what actually forces it is power. the injector is putting 0.5 mW into the fiber. spontaneous Raman scattering off that pump spreads a noise floor across a wide swath of spectrum around it, and 0.5 mW against a detector counting single photons is not a contest. i went through the arithmetic of this in the coexistence post; the summary is that you can run a classical signal alongside a quantum channel in the same fiber, but only after you have thought hard about launch power and spectral distance, and half a milliwatt one channel over is nowhere near that regime.
which reframes the problem. the time multiplexing is a power problem wearing a wavelength problem’s clothes. the fix is not a faster switch, it is a dimmer pilot. and that work exists: continuous stabilization from a dim coexisting reference does the measurement with heterodyne detection instead of a polarimeter, which buys back enough sensitivity to run the reference at a power the quantum channel can tolerate. put that on this link and the 3 second window and the 55 second timeout both stop existing as concepts.
one honest counterargument, which is the reason we did not do that here. the fidelity curve in Fig. 3(b) is a snapshot. wavelength-dependent birefringence is itself a property of the fiber that drifts, so the offset between what the pilot experiences and what the photons experience is not a fixed number you can calibrate away once. it wanders too, on the same aerial fiber, for the same reasons. at zero detuning that entire class of error is identically zero, and on a link this unstable that is worth paying for. how much it actually wanders on aerial fiber is, as far as i know, not measured anywhere, and somebody should measure it.
the duty cycle
so: pilot and photons share a channel, therefore they cannot share the fiber at the same time, therefore a pair of optical switches alternates the link between compensating and distributing.
that gives the link a duty cycle, and the duty cycle is the headline number. over 24 hours the compensation sessions consumed 7.2% of the link, leaving 92.8% uptime for entanglement.
the most useful thing about that duty cycle is not in the paper, and it took me a while after publication to see it. start from what is published. the window is 3 seconds and the uptime is 92.8%. so
3 / (3 + t_avg) = 0.928 → t_avg ≈ 233 ms
the average compensation session is a shade under a quarter of a second. that is a fine number. it is also completely unrepresentative, because Fig. 5(b) shows the distribution: most sessions sit on a floor around 100 ms, and a handful run to tens of seconds, with one at 41 s.
so take 100 ms as the normal case and ask what fraction of sessions have to be timing out at 55 s to drag the mean up to 233 ms. the answer is about a quarter of one percent. over 24 hours the link runs roughly 26,700 cycles, so that is about 64 events in a day, and those 64 events account for around 57% of all the downtime.
that changes what the fix looks like. the paper’s outlook section proposes two things: faster hardware, meaning lithium niobate polarization controllers instead of the current ones, and a different algorithm, specifically a reverse-operator scheme that computes the inverse transformation from a fixed number of measurements instead of descending toward it.
the first one improves the average, and the average is not the problem. cut the ordinary session by a factor of ten, from 100 ms to 10 ms, and uptime goes from 92.8% to 95.5%. worth having. but every bit of the remaining 4.5% is then timeouts, and you have spent a hardware upgrade to get there. go the other way, leave the controller alone and stop the sessions timing out, and you get 96.8%. do both and you get 99.7%.
the second one attacks the tail, and the tail is the entire problem. a gradient descent on a target that is moving faster than the loop can chase it does not converge slowly, it fails to converge, and then it times out. a scheme with a fixed measurement count cannot fail that way. it has a worst case equal to its best case. that is the property that matters here, and it is worth saying more loudly than we said it.
there is a preprint from a few weeks ago claiming exactly this direction, exponential speedup of polarization stabilization for long distance DWDM quantum networks. i have not read it carefully enough to summarize honestly, but the title is aimed at the right thing.
one more number while we are here. the 92.8% is link uptime, and inside each 3 second window the time tagger only runs for 2 seconds, because you have to stay clear of the next session. so the measurement duty cycle is 2 / 3.233, or about 62%. if you were quoting a service level to somebody, 62% is the honest figure and 92.8% is the one about the fiber.
what came out
with the stabilization running:
- roughly 1500 entangled pairs per second detected across the link, in a 1.6 ns coincidence window, with the polarization analyzers out of the path.
- coincidence-to-accidental ratio around 12.
- a 24 hour time-averaged CHSH parameter of S = 2.34 ± 0.37, against a classical bound of 2, with the violation holding for more than 20 of those hours.
and the control, which is the part i find more persuasive than the result: with the stabilization off, the reference fidelity collapsed within minutes and S fell below 2 after about three hours and never came back for the remaining thirteen.
there is one more number in there that i want to pull out, because it is the best evidence in the paper for something the paper does not quite say. throw away the data taken immediately after a timed-out session, and the 24 hour average becomes S = 2.39 ± 0.14. the mean barely moves, 2.34 to 2.39. the error bar drops by a factor of 2.6. the timeouts were not degrading the average result. they were producing the variance.
| NIST analyzer basis | 0° (H) | 45° (D) | 90° (V) | 135° (A) |
|---|---|---|---|---|
| Visibility, at the source | 0.97 | 0.96 | 0.97 | 0.96 |
| Visibility, after 62 km | 0.87 | 0.78 | 0.76 | 0.79 |
| S, at the source | 2.69 ± 0.02 | |||
| S, after 62 km | 2.26 ± 0.04 | |||
the state went in at S = 2.69 and came out at 2.26. the interesting question is where the 0.43 went, and the answer is not the one you would expect from a paper about polarization.
where the Bell violation actually leaked out
three things happen to the state between the source and the far analyzer, and they are worth separating because only one of them is about polarization.
accidentals. a coincidence window is a bucket. real pairs land in it, and so does anything uncorrelated that happens to arrive at the right moment. the bucket here is 1.6 ns wide, the CAR is about 12, and for a fringe sitting on a flat accidental background the visibility ceiling is
V = CAR / (CAR + 2) = 12 / 14 = 0.857
that alone takes 0.965 down to about 0.83, which is most of the loss.
leftover polarization drift. the compensator hands back a fiber at better than 99% fidelity and then the fiber starts moving again immediately, for the whole 3 seconds you are using it. that is a real effect and it is the one the experiment is nominally about. it is also, on this budget, the smaller term.
the coincidence window itself. that 1.6 ns is not set by anything quantum. the SNSPDs jitter by about 100 ps. chromatic dispersion over 62 km smears the idler by about 300 ps. those add to something like 320 ps. the remaining ~1.5 ns, which is to say nearly all of it, comes from the electrical-optical-electrical link: the NIST detector clicks get turned into laser pulses, sent to UMD down a parallel fiber, and turned back into electrical pulses, so that one time tagger can see both ends of the experiment.
which leads to the sentence i would have argued for putting in the abstract: the largest single lever on this link’s Bell violation is a classical timing problem.
you can see the shape of it by closing the window on the link as it is. narrower window, fewer accidentals, better S, and fewer pairs per second, because you are now cutting into the peak. that is a trade, and 1.6 ns is roughly where we stopped trading.
narrowing the peak is the same move without the bill. shrink the jitter from 1.6 ns to the ~320 ps the detectors and the dispersion would give you on their own, and the window shrinks with it while staying matched to the peak. real coincidences are untouched. accidentals fall by a factor of five, because they scale with window width and nothing else. CAR goes from 12 to about 60, and S goes from 2.26 to something around 2.55.
that is a bigger improvement than anything the polarization stabilization could deliver, and it comes from replacing an E-O-E timing hack with a real synchronization scheme. which is, satisfyingly, the exact subject of the coexistence post and of the DC-QNet clock characterization work done down the hall. the pieces are in the same building. they were not in the same experiment.
in fairness to the E-O-E link: the two fibers are in the same bundle from the same TX/RX pair, and the relative delay between them moved only about 110 ps over the full 24 hours. as a stability solution it is excellent. it is the absolute jitter, not the drift, that costs.
the 41 second outlier
there is a datapoint in Fig. 5(a) that sits visibly off its fringe, and Fig. 5(b) tells you why: the compensation session immediately before it ran for 41 seconds. excluding it lifts that basis from V = 0.76 to 0.82 and the whole measurement from S = 2.26 to 2.30.
we reported it both ways, which is the right thing to do, and the paper explains it as a temporary volatile change in the fiber environment. from inside, the honest version is: something happened to that cable and we do not know what. wind, a truck, someone’s boot on a strand-mounted splice case. the same story explains the timeout clusters in Fig. 6(b), which fall in the middle of the day and around the evening rush and not at three in the morning.
this is the same story as the ±0.37 error bar above, at the scale of a single datapoint. the general form of it, and the reason it is worth more than an anecdote, is that a 41 second compensation session is not just 41 seconds of lost link. it is 41 seconds of the fiber moving while the controller is failing to catch it, and when the session finally times out it hands back a channel that is still moving. so the next 3 second window is bad too. the damage outlives the outage. that is why timeouts show up in the S-value trace and not only in the uptime column, and it is another argument for an algorithm that cannot fail to converge.
what i’d push on next
1. dim the pilot and stop time multiplexing. the polarization penalty for moving one ITU channel away is about 0.01 in S. the barrier is the 0.5 mW launch, not the 0.8 nm. heterodyne detection of a dim reference is already demonstrated. putting it on this link deletes the duty cycle, the 55 s timeout and the 3 s window in one move, and turns a threshold machine into a continuous servo.
2. measure how much the wavelength offset itself drifts. this is the honest objection to point 1 and nobody has the number. Fig. 3(b) is one snapshot on one afternoon. if the pilot-to-photon transformation offset on aerial fiber is stable to a few tenths of a percent over hours, point 1 is free. if it wanders, point 1 needs a slow outer calibration loop and somebody should design it.
3. fix the clock, not the controller. 1.6 ns of coincidence window is buying five times more accidentals than the physics requires. an actual synchronization scheme in place of the E-O-E link is worth roughly 0.3 in S, which is more than doubling the source brightness would give you.
4. report the tail, not the mean. 7.2% average downtime is the number that makes the abstract, and the 0.37 error bar next to the headline S is what the tail did to it. the distribution behind it is 64 pathological events and 26,000 uneventful ones. any operator running a quantum link is going to care about the 64, because that is what a service level agreement gets written against. the field should be reporting compensation-time percentiles the way we report latency percentiles, and i would like p99 and p99.9 to become normal in these papers.
5. put an admission-control decision on top of it. this connects to the thing i keep saying about flex-grid allocation. if a controller can tell, from the reference fidelity trace alone, that the fiber has entered a regime where sessions are going to time out, the right response is not to keep trying. it is to tell the application that this link is unavailable for the next few minutes and let it fail over or wait. right now the link has exactly one behavior under stress, which is to keep grinding, and the cost of that shows up as bad data rather than as an honest outage.
6. someone should do this on a worse link. 70% aerial and 18 dB is a hard link. it is not the hardest one in anybody’s metro plant. the interesting experiment is the one that finds where this approach actually breaks, and we did not find that here, because it did not break.
the short version
- polarization entanglement is the easy kind to make and the fragile kind to move. buried fiber holds its polarization transformation for hours. aerial fiber holds it for seconds.
- the link is 62 km NIST to UMD, about 70% strung between poles, 18 dB end to end, of which 5.6 dB is connectors and splices rather than glass.
- unstabilized, the transmitted state falls below 95% fidelity in under 20 seconds during the day, and S drops through the classical bound after about 3 hours and stays there.
- the fix is a bright polarization reference in the same channel as the photons, six reference states, gradient descent on the worst of their fidelities, 98% to trigger and 99% to stop.
- same channel means no coexistence, which means time multiplexing: 3 s of entanglement, then a check, forever.
- result: about 1500 pairs/s, CAR 12, S = 2.34 ± 0.37 averaged over 24 hours, violation held for more than 20 of them, 92.8% uptime.
- drop the data taken right after a timed-out session and it becomes S = 2.39 ± 0.14. the mean barely moves and the error bar shrinks by 2.6x. the timeouts were making the variance, not the average.
- the honest uptime for measurement is 2 s in every 3.23, about 62%. the 92.8% is a fact about the fiber, not about the service.
- back out the numbers and the average compensation session is 233 ms, of which the typical case is ~100 ms. about 64 timeout events in a day cause roughly 57% of the downtime. the mean is not the problem, the tail is. ten times faster hardware buys 95.5% uptime; not timing out buys 96.8%; both together buy 99.7%.
- so faster hardware buys almost nothing, and an algorithm that cannot fail to converge buys everything. the paper says both; only the second one matters.
- the biggest available gain is not quantum. the 1.6 ns coincidence window is set by jitter in a classical timing link, not by the detectors or the dispersion. fixing it is worth about 0.3 in S, more than any plausible improvement to the stabilizer.
- the pilot could probably sit an ITU channel away for a penalty of about 0.01 in S. what keeps it in the same channel is 0.5 mW of launch power, and that is a solvable problem with an existing solution.
what i took away from a year of this is not really a quantum lesson. most of the hard part was ordinary network engineering in an unusual costume: a duty cycle, a tail latency problem, and a clock distribution hack that turned out to set the error budget. the entangled photons mostly did what they were supposed to. it was the cable in the wind and the timing link we bolted onto it that decided how the numbers came out, and i suspect that will keep being true for a while.