Cosmic Rays and Silent Data Corruption in Data Centres

Cosmic-ray neutrons are real, but field studies point to hardware defects behind most DRAM faults and the silent core errors operators have traced.

15 min read

Key takeaways

  • Cosmic-ray neutrons are real: about 13 per square centimetre per hour above 10 MeV at New York sea level, as quoted by Asorey and Mayo-Garcia 2022. Particle strikes, mainly cosmic-ray neutrons, cause most faults in on-chip SRAM, the fast memory inside processors (Sridharan et al. 2015; Baumann 2005).
  • In DRAM, the main memory, field studies point mostly to hardware faults. Google concluded that errors are “unlikely to be dominated by soft errors” (Schroeder et al. 2009). On NERSC’s Hopper system at most 44.5% of DRAM faults were transient (our sum of the paper’s Table 2), and transient does not mean radiation.
  • Silent data corruption in compute cores, as the operators have traced it, comes from defective chips. Alibaba found 3.61 faulty CPUs per 10,000 (Wang et al. 2023). Google reports a few per several thousand machines (Hochschild et al. 2021). Meta says its rate is “much higher than cosmic-ray-induced soft errors” (Meta 2025).
  • Altitude matters on a mountain, not across European hubs. Los Alamos’s Cielo system is estimated to receive 5.53 times New York’s neutron flux. Our calculation with the Ziegler formula gives roughly 1.0 to 1.7 times between Amsterdam and Madrid.
  • ECC memory, scrubbing and fleet screening beat shielding. ECC covers memory and cache bit flips. Screening covers the defective cores that ECC cannot see. A foot of concrete cuts the high-energy flux only about 1.4 times (our calculation).

Most memory faults, and the silent compute-core corruption that large operators have traced, come from defective or worn silicon in the studies we read, not from cosmic rays. Radiation does flip bits, and it dominates faults in processor SRAM. In the DRAM and CPU fleet studies we can read, though, hardware defects explain more of what operators see. The practical answer is to buy ECC everywhere, track faults per DIMM, and screen every processor for silent miscomputation. Shielding and site selection come far down the list. The reference table below sets out each cause, its evidence and the action it calls for.

How a neutron from space flips a bit

The story starts with a particle that is not a neutron. Primary cosmic rays arrive from space, and to reach sea level they need an energy above about 1 GeV. They hit nuclei in the upper atmosphere and start a shower of secondary particles. Near the ground, about 95% of the strongly interacting particles left are neutrons, according to Ziegler (IBM, 1996). That paper quotes 97% in its body text, so read it as 95 to 97%.

The flux is about 13 neutrons per square centimetre per hour above 10 MeV at New York sea level. That figure is quoted by Asorey and Mayo-Garcia (citing Gordon et al. 2004). Texas Instruments quotes about 13 per square centimetre per hour at sea level. Almost every neutron passes through a chip untouched. Ziegler estimated in 1996 that one in 40,000 interacts within 10 micrometres of a circuit. That gives about 2.5 hits a year for each square centimetre of active circuitry. It is an old, illustrative estimate.

A hit matters because of what it leaves behind. The neutron knocks a silicon nucleus out of place, and the fragment ionises the material around it. Our pieces on the Landau distribution and the straggling function explain how a charged particle deposits charge in a silicon layer. If a circuit node collects more charge than its critical charge, the stored value flips. Baumann (Texas Instruments, 2005) says critical charge depends on node capacitance, operating voltage and the strength of the feedback transistors. Lower voltage means a smaller critical charge.

A single strike can sometimes upset more than one cell. Ziegler reported double soft fails from one energetic particle, shown on DRAMs with a proton beam and at Denver. We found no field rate for these multi-bit upsets in modern parts. Some error-correcting designs interleave bits to protect against multi-bit faults (Sridharan et al. 2015).

Two other sources add to the neutron story. Thermal neutrons are slow neutrons. When one is captured by boron-10, the nucleus splits and releases an alpha particle and a lithium ion. Boron in a glass layer called BPSG was the main source of soft errors in some 0.25 and 0.18 micrometre SRAM, and Baumann notes that makers removed it from virtually all advanced processes. Alpha particles from uranium and thorium traces in packaging were the main cause of DRAM soft errors in the late 1970s. Purer materials fixed that. Both are properties of how a part was made, and an operator cannot change them.

The scaling trend is not what many readers expect. Baumann reports that DRAM soft-error rate per bit fell about 4 to 5 times per generation, more than 1,000 times over seven generations. The DRAM rate per system stayed almost flat because systems hold more bits. SRAM per-bit rates rose in early generations, then saturated below about 250 nm. His worked example is a processor with a lot of embedded SRAM: more than 50,000 FIT per chip, or about one soft failure every two years. A system of 100 such chips would see one about once a week. He concludes that “error correction is mandatory”. (FIT is failures per billion device hours.)

What field data says about memory

Google’s study of its fleet, DRAM errors in the wild (2009), covered 2006 to 2008 hardware. More than 8% of DIMMs had errors in a year. About a third of machines saw at least one memory error a year, and 1.3% had an uncorrectable error. The authors concluded that errors are dominated by hard errors, meaning repeating faults in the hardware. They add a limit: their data cannot reliably separate hard from soft errors, so this is an inference from the way error rates track machine utilisation.

The Los Alamos and NERSC study (Sridharan et al. 2015) counted faults instead of errors, which is the better unit. A single stuck cell can produce thousands of error messages. On the Hopper system in Oakland, about 78.9% of DRAM faults were single-bit. Adding the transient rows of the paper’s Table 2 gives 44.5% of faults as transient. The permanent rows give 55.4% (our sums; rounding leaves 99.9%). Transient faults can have causes other than radiation, so 44.5% is a ceiling for the neutron share on that machine. The paper says there are “clearly many other causes of faults”.

The altitude comparison is the most instructive part. Cielo, at about 7,300 feet in Los Alamos, is estimated to receive 5.53 times New York’s neutron flux. Hopper, at 43 feet in Oakland, is estimated to receive 0.91 times. That is a flux ratio of 6.08 (our division). If DRAM single-bit transient faults followed the flux, Cielo’s rate per device would be about six times higher. The paper’s Table 3 gives Cielo-to-Hopper ratios of 1.70, 1.61 and 0.77 for single-bit transient faults for three vendors (our division). So for these DDR3 parts only part of the transient faults tracked neutron flux, and the memory vendor mattered a lot. The authors write that “judicious choice of the memory vendor can reduce this impact”. The Cielo numbers come from an earlier paper with a different observation period, so treat the ratios as indicative.

Processor SRAM behaves differently. Cielo had substantially more corrected SRAM faults than Hopper, but its uncorrected L2 and L3 cache errors were not substantially higher, which the authors attribute to ECC and bit interleaving in those caches. The authors wrote that “an appropriately-designed system need not be less reliable when located at higher elevations.” Without that protection the result can be bad. On the ASC Q supercomputer at Los Alamos (2003), one SRAM structure had parity but no ECC. Its parity errors caused node crashes and made up about half of all failures, and beam tests supported a neutron explanation, according to a 2010 summary by Michalak (LANL slides). Modern server caches carry ECC, so this is a lesson about unprotected arrays.

Two cautions on what to measure and buy. First, raw error counts can mislead. In the Los Alamos paper Hopper logged four times Cielo’s memory error rate but only 0.625 times its fault rate. Track faults per DIMM. Second, plain SEC-DED ECC (single-error correct, double-error detect) can miss faults. The authors estimate that, on Cielo’s DDR3 parts, faults SEC-DED cannot detect reach over 21 FIT per DRAM device for the worst of three vendors (1.8 and 0.2 for the others), and call SEC-DED poorly suited to modern DRAM (submitted version). Google reported that chipkill-correct ECC, which corrects any error confined to one DRAM device, cut uncorrectable errors 4 to 10 times against SEC-DED across Google platforms (Schroeder et al. 2009; Sridharan et al. 2015, section 8.1). Both results predate DDR5 with on-die ECC, so check what your platform actually provides.

Silent corruption in the cores is a defect problem

ECC protects stored bits. It does nothing when a core computes the wrong answer, which is the case that operators call silent data corruption (SDC). Google’s Cores that don’t count (2021) describes “mercurial cores” that miscompute without raising any error. It reports on the order of a few per several thousand machines. Typically one core in a multicore chip fails. Failures usually repeat, often get worse with time, and depend on frequency, voltage and temperature. The authors say SDCs were long ascribed to random causes such as alpha particles and cosmic rays. They view them instead as symptoms, with high-rate faulty cores as a new cause.

Meta reaches the same conclusion. In Silent Data Corruptions at Scale (2021) it says SDCs are not limited to soft errors due to radiation and environmental effects, and that they are repeatable at scale. A later paper, Detecting silent data corruptions in the wild (2022), puts the rate at one in a thousand silicon devices and says this is not limited to particle effects or cosmic rays. Treat “one in a thousand” as an order of magnitude: the paper does not define the denominator. Meta’s 2025 engineering post says the rate is “much higher than cosmic-ray-induced soft errors”. Treat the comparison with care: one-in-a-million is a modelled soft-error rate and one-in-a-thousand is Meta’s reported fleet figure, whose denominator is not defined.

The only fleet measurement with a published denominator is from Alibaba Cloud (Wang et al., SOSP 2023). Over more than one million CPUs in 28 data centres and 32 months of testing, 3.61 per ten thousand CPUs were found to cause SDCs. The paper’s unit is per ten thousand, not per cent. That is 0.0361%, or about 1 in 2,770 (our conversion). The parts came from one manufacturer. Of the faulty CPUs, 90.36% were found in pre-production tests.

The faults did not look like radiation upsets. Some SDCs were highly reproducible, while others appeared only under specific conditions. Frequencies ran from 0.01 times a minute to hundreds of times a minute. In some settings the frequency grew exponentially with temperature, with triggers inside the normal operating range. Most corrupted results had one flipped bit, but a considerable number had two or more, and 51.08% of bit flips went from zero to one. The authors point out that models built on radiation assume independent, identically distributed bit flips, and say their observations challenge this.

AI fleets show a similar picture, with a caveat. Meta’s Llama 3 training run on 16,384 H100 GPUs logged 419 unexpected interruptions in 54 days (Grattafiori et al. 2024, Table 5). About 78% were attributed to confirmed or suspected hardware issues. Silent data corruption was listed for 6 of the 419, or 1.4%. Meta does not attribute the GPU memory or SRAM entries to radiation. An SDC is by definition hard to detect, so 6 is a lower bound. ByteDance reports that NVIDIA’s diagnostic tool reached only 70% recall in its production environment (Wan et al. 2025). Its listed causes include numerical instabilities, race conditions and thermal variations, and it does not discuss cosmic rays.

Altitude: a calculation for European sites

Ziegler gives the atmospheric depth A in grams per square centimetre as A = 1033 – 0.03648 H + 0.000000426 H squared, with H the elevation in feet. The neutron flux relative to sea level is exp((1033 – A) / 148), where 148 g/cm2 is the neutron absorption length. His own check is Denver at 5,280 feet: 3.4 times sea level. We applied this to round elevations. They are our approximations, not surveyed campus heights, so check your own site. The formula ignores the geomagnetic field and weather.

Site (approximate elevation)Elevation usedFlux relative to sea level
Amsterdam0 m1.00
Frankfurt110 m1.09
Zurich410 m1.39
Madrid650 m1.67
A mountain site1,500 m3.14
Los Alamos (7,300 ft)2,225 m5.19 (the paper reports 5.53)

The formula lands about 6% below the Cielo figure from Sridharan et al. Asorey and Mayo-Garcia simulated high-energy flux (above 50 MeV) for supercomputing centres. Their Madrid-area site gives 1.71 times Jülich and their Munich site 1.42 times. The formula gives 1.60 and 1.34 for the same elevations (700 m and 471 m against 100 m). These are simulations, not measurements. Across European hubs the altitude spread is therefore about 1.0 to 1.7 times, which is small next to the 5 times at Los Alamos. The same paper reports that ordinary pressure changes move the flux by a few percent, and seasonal air-density changes by up to about 15% in one energy bin at Los Alamos. A site at 1,500 m is a different case, at roughly three times sea level by the formula.

Shielding does not close the gap. Ziegler measured a neutron attenuation length in concrete of 216 g/cm2. Concrete is 2.45 g/cm3, so a 30 cm slab is 73.5 g/cm2 (a full foot, 30.5 cm, is 74.7) and cuts the flux by about 1.4 times (our calculation). About 3 m of concrete would give roughly 30 times. A data-hall floor changes nothing that matters. Concrete and cooling water can also raise the thermal-neutron flux. In Oliveira et al. (2020), concrete raised it by up to 20% and rain by up to 2 times. In beam tests of older accelerators, thermal neutrons accounted for 4.2% of a Xeon Phi’s SDC rate in New York and 29% of a K20 GPU’s SDC rate at Leadville, and 39% of an AMD APU’s DUEs at Leadville (Oliveira et al. 2020, assuming a concrete slab). We found no thermal-neutron data for current GPUs.

Get the next one by email

Power, cooling, chips and the rules that shape them. No spam, unsubscribe any time.

Reference table: cause, evidence and action

Read the evidence column with its year. Several sources describe older hardware, and the last column is our reading of what each source supports.

CauseEvidence (source, year)What it affectsOperator action
Neutron and alpha upsets in DRAMHard faults dominate Google’s fleet (2009). At most 44.5% of Hopper DRAM faults were transient (sum of Table 2, 2015). Vendor effect strong.Main memory (DIMMs)Use ECC DIMMs everywhere. Prefer chipkill-class protection to plain SEC-DED. Enable patrol scrubbing and command/address parity. Track faults per DIMM, not raw error counts.
Neutron upsets in on-chip SRAMMost SRAM faults in the field are particle-induced transients (2015). An SRAM structure (BTAG) with parity but no ECC caused node crashes on ASC Q; the neutron explanation was a hypothesis consistent with beam tests (2010 slides).Caches, registers, buffersChoose processors and GPUs with ECC on caches and register files. Rely on silicon design, not site choice.
Altitude5.53 times New York’s flux at Los Alamos (2015). 1.0 to 1.7 times across Amsterdam to Madrid (our calculation from Ziegler 1996, supported by a 2022 simulation).Rate of all neutron upsetsTreat altitude as a minor criterion for European hubs. At 1,500 m or more, require full memory and cache ECC and log corrected errors.
Thermal neutronsThermal share of FIT ranged from 4.2% to 39% depending on device, site and error type (Oliveira 2020). Higher near concrete and in rain.Mostly GPUs and acceleratorsCheck device ECC. Ask vendors about boron-10 content for critical parts. Do not redesign a hall for it.
ShieldingAbout 1.4 times attenuation per 30 cm of concrete (our calculation from 1996 data).High-energy neutron fluxDo not specify. ECC and screening cost less and cover more.
Defective or worn coresA few per several thousand machines (Google 2021). 3.61 per 10,000 CPUs (Alibaba 2023). About one in a thousand devices (Meta 2022).Computation: wrong results with no errorScreen at receiving. Run scheduled test suites and in-production tests. Quarantine a core on mismatch. Control core temperature and workload.
Silent errors in AI training6 of 419 Llama 3 interruptions were SDC (2024). A vendor tool had 70% recall at ByteDance (2025).GPU training jobsAdd checkpoints and verification such as loss-spike and gradient checks. Run vendor diagnostics plus your own tests. Keep spare nodes.
Corruption anywhere in the data pathMeta traced a file-size computation returning zero in a decompression pipeline, with clean system logs (2021).Stored and transferred dataUse end-to-end checksums and application-level validation. Avoid single-copy computation for critical outputs.

What to do

Start with the controls that pay off whatever the cause. Buy ECC memory with chipkill-class protection, turn on scrubbing, and log faults per DIMM. If you are sourcing memory in the current market, our report on the 2026 data-centre memory shortage covers supply. The evidence above argues for holding the ECC requirement even then.

Then screen the processors. Alibaba found 90.36% of its faulty CPUs in pre-production testing. Meta’s out-of-production scanner reaches full fleet coverage in five to six months. Its in-production tester covered 70% of the same detections within 15 days. Those coverage figures are relative to one defect family, not to all defects. Test again later, because in one Meta experiment a device passed a daily repeat of the same computation for six months and then failed.

Last, put altitude and shielding in their place. Use altitude as a tiebreaker for a high site, and do not pay for concrete.

What we don’t know

No source we read measures the soft-error rate at a European data centre. Every site comparison above is modelled. The sea-level flux figure rests on secondary citations. No field rate for multi-bit upsets in DDR5 or HBM is available. The thermal-neutron data are from older devices. Google’s DRAM result is an inference, as the authors say. Alibaba’s population is one manufacturer, and its test suite can miss faults. Meta’s “one in a thousand” has no stated definition. Radiation may matter more in parts that no study here covers, such as current GPU memory, and this article cannot rule that out.

Sources

Get the weekly brief

Sourced analysis of data-centre engineering and regulation, every Monday. No spam, unsubscribe any time.

Alex Turner
Alex Turner

Alex Turner is an engineer who works on data centre infrastructure. He writes about data centres as physical buildings, where the choices are forced by heat and by the grid long before anyone decides them.

Articles: 8

Leave a Reply

Your email address will not be published. Required fields are marked *