Quick answer

ECC does not stop a DIMM failing; it warns you first. Google’s fleet study found 70 to 80 percent of uncorrectable errors were preceded by a correctable error that month or the one before. Buy the ASRock Rack B650D4U at $299 and confirm EDAC reports SECDED.

On this page

Google logged memory errors across its server fleet for two and a half years. Uncorrectable errors — the kind that take a machine down — hit 1.3% of machines per year. Correctable errors hit about a third of them. And in 70 to 80% of the uncorrectable cases, the same machine had already logged a correctable error that month or the month before.

That last number is the argument for ECC on a home server, and it is not the argument people make. ECC’s value is not that it repairs the bit. It is that it names the failing DIMM weeks early, in a counter you can read.

Which means ECC you cannot read is ECC you did not buy. Most of this page is how to prove yours is switched on.

Some retailer links below are affiliate links; if you buy through them TechFuel HQ may earn a small commission at no extra cost to you. As an Amazon Associate I earn from qualifying purchases. Commissions never influence what gets recommended — see our disclosure.

What the two fleet studies actually found

There are two canonical field studies, and between them they cover most of what is publicly known about how real DRAM fails in real racks.

Google fleet (2006–2008)Facebook fleet (2013–2014)
Scale~2.5 years, millions of DIMM days14 months, billions of device days
Servers with a correctable error, per year32.2%9.62%
Uncorrectable errors1.3% of machines per year0.03% of servers per month
Median errors on an affected machine25–611 per year by platform≤9 per month
Where the errors concentratehighly skewed by DIMMtop 1% of servers hold 97.8%

Schroeder, Pinheiro and Weber reported FIT rates of 25,000 to 70,000 per Mbit against the 200 to 5,000 that lab work had predicted, and wrote that memory errors “are not rare events.”

Meza, Wu, Kumar and Mutlu found a much lower yearly incidence seven years later. The two studies disagree by roughly 3x on how many servers see an error, and the Facebook paper says why in its own text: device generations improved over the better part of a decade, and its 9.62% figure lines up with the 5.48% to 9.10% another recent study had measured. The older number describes DDR1 and DDR2 era hardware; the newer one describes DDR3.

They also part company on error character. Google reported strong evidence that errors are dominated by hard errors — repeatable physical defects — while noting its collection could not distinguish hard from soft directly. Facebook, after separately accounting for memory controller and channel failures, found that one-off spurious failures showed up on the largest share of affected servers, 56.03%. These are different measurements of different things, and the practical conclusion survives both: some fraction of your DIMMs will start throwing errors, and the machines that do it are wildly unrepresentative of the fleet.

That skew is the part that matters at home. You do not own a fleet. You own two to four DIMMs. On the Facebook numbers the median affected server logged nine correctable errors a month while the mean was 497 — the average is 55 times the middle, because a handful of machines were on fire. You are either fine or you are the fire, and without ECC there is no way to know which.

What ECC is protecting, and what it is not

Side-band ECC — the real thing, the kind in server RAM — stores check bits in an extra DRAM chip and widens the memory bus to carry them. The memory controller generates the code on write and verifies it on read. The common scheme is SECDED: correct any single-bit error, detect any double-bit error. That is a whole-path guarantee from the controller through the connector to the cells and back.

Uncorrectable errors are the failure this prevents. In Google’s environment a single uncorrectable error was treated as serious enough to shut the machine down and replace the DIMM. On your NAS the same event is a kernel panic mid-scrub, or a silent bad byte written into a snapshot you will keep for three years.

But the correction is the smaller half of what you are buying. The bigger half is the log line.

DDR5 on-die ECC is not ECC, and the listings will not tell you

Every DDR5 chip has on-die ECC. It is a manufacturing yield measure, not a reliability feature you own, and PassMark’s MemTest86 documentation is blunt about the limits: on-die ECC “does not provide end-to-end protection,” it “does not detect or prevent errors that occur during transmission between the memory controller and the memory module,” and it “is completely invisible to the CPU and memory controller,” providing no feedback to the rest of the system.

Read that last clause again against the reason you wanted ECC. On-die ECC can never tell you a DIMM is degrading, because it never tells the operating system anything at all.

Retailers put the phrase in titles anyway. A live Newegg listing for a Kingston FURY Beast 32GB DDR5-5600 gaming module — a black-heatspreader desktop stick — describes itself in its own title as “On-die ECC Unbuffered 288Pin UDIMM Standard Desktop Computer Memory.” It is not ECC memory. It will never populate an EDAC counter. It carried $709.99 in the August 17 feed, which is server-memory money for a gaming stick.

So buy by part number, not by keyword. Real ECC UDIMMs are sold under server part numbers — Kingston KSM, Micron MTC with an EC in the code, Samsung M324R — and those are the strings that appear on a motherboard’s memory QVL. The word ECC in a product title, on its own, tells you nothing at all.

What has to be true before any of this works

Four things, and all four are decided before you power on. If you are still choosing the platform, the Proxmox hardware requirements guide covers the rest of the spec around this decision.

The CPU. AMD’s own product page for the Ryzen 9 9950X lists ECC Support as “Yes (Requires mobo support).” That parenthetical is doing all the work. AMD ships the memory controller capability across the desktop line and then hands responsibility for enabling it to whoever made your board. Intel splits it by platform instead: ECC on Core processors needs a workstation chipset, which is why Supermicro’s X13SAE-F pairs W680 with 12th–14th Gen Core and lists “192GB Unbuffered ECC/non-ECC UDIMM.”

The board. This is where consumer boards fail you quietly. A spec line reading “supports DDR5 ECC/non-ECC UDIMM” tells you the slot accepts the module. It does not tell you the firmware turns correction on. Server boards publish a memory QVL naming exact ECC part numbers; ASRock Rack puts one behind the Memory QVL tab on the B650D4U’s product page, and that list, not the spec bullet, is what to buy from. A named part on a published list is a commitment. “ECC compatible” in a marketing bullet is not.

The RAM. Unbuffered ECC (ECC UDIMM) for these platforms, not registered (RDIMM), which needs a server socket. In the August 17 Newegg feed the cheapest genuine 32GB DDR5 ECC UDIMM row sits at $849.99 against $499.00 for the cheapest non-ECC 32GB module — roughly $350 a stick. Both numbers are inflated by the 2026 DRAM crunch and both are dated. The gap is the shape of the decision.

The firmware setting. MemTest86 documents the option names you are looking for: DRAM ECC Enable on AMI, ASUS, ASRock and MSI firmware, or ECC Mode on some ASUS boards.

Then verify: four independent things have to line up, and nothing on the POST screen confirms they did.

Verify it: four commands

1. Ask the firmware what it thinks

sudo dmidecode -t 16
Physical Memory Array
        Location: System Board Or Motherboard
        Use: System Memory
        Error Correction Type: Multi-bit ECC
        Maximum Capacity: 128 GB
        Number Of Devices: 4

Error Correction Type is dmidecode printing the SMBIOS Type 16 field. DSP0134 v3.7.0 defines exactly seven values for it: Other, Unknown, None, Parity, Single-bit ECC, Multi-bit ECC, CRC. None is a clear negative. Anything with ECC in it is a claim.

Then check the modules themselves:

sudo dmidecode -t 17 | grep -E 'Locator:|Total Width|Data Width|Size:'
        Total Width: 72 bits
        Data Width: 64 bits
        Size: 32 GB
        Locator: DIMM A1

The SMBIOS spec defines Total Width as the width “including any check or error-correction bits,” and says that “if there are no error-correction bits, this value should be equal to Data Width.” Equal widths mean the module carries no check bits. Wider than the data width means it does — 72 against 64 on an ECC UDIMM, wider still on a registered module. This is the one dmidecode check that describes physical hardware rather than a firmware opinion.

And that is the limit of what dmidecode is worth. Its own manual page ends with a BUGS section that reads, in full: “More often than not, information contained in the DMI tables is inaccurate, incomplete or simply wrong.” The firmware can advertise Multi-bit ECC on a board that never enabled it. Treat this step as a screen, not a verdict.

2. Ask the memory controller

grep -H . /sys/devices/system/edac/mc/mc*/dimm*/dimm_edac_mode
/sys/devices/system/edac/mc/mc0/dimm0/dimm_edac_mode:SECDED
/sys/devices/system/edac/mc/mc0/dimm2/dimm_edac_mode:SECDED

This is the check that settles it. EDAC is the kernel’s Error Detection And Correction subsystem, and dimm_edac_mode reports, in the ABI documentation’s words, “what type of Error detection and correction is being utilized.” The kernel picks from a fixed string table in edac_mc_sysfs.c: Unknown, None, Reserved, PARITY, EC, SECDED, S2ECD2ED, S4ECD4ED, S8ECD8ED, S16ECD16ED.

  • SECDED or an S*ECD*ED chipkill mode: correction is running. You are done.
  • None: the controller is not doing ECC, whatever the DIMMs are.
  • Nothing printed, or /sys/devices/system/edac/mc/ empty: no EDAC driver bound. Load the one for your platform — amd64_edac for AMD, ie31200_edac or igen6_edac for Intel client parts, skx_edac or i10nm_edac for Xeon Scalable. An empty directory is a missing driver, not a verdict on your hardware.

3. Read the counters, and record the baseline

grep -H . /sys/devices/system/edac/mc/mc*/{ce_count,ue_count,ce_noinfo_count,ue_noinfo_count}
grep -H . /sys/devices/system/edac/mc/mc*/dimm*/dimm_{label,ce_count,ue_count}

The ABI’s note on ce_count is the operational instruction: “This count is very important to examine. CEs provide early indications that a DIMM is beginning to fail. This count field should be monitored for non-zero values.” ce_noinfo_count carries corrected errors the controller could not attribute to a slot — the docs describe that state as memory being “handicapped, but operational” — and it needs watching too.

Write down what these read on a healthy boot. A number means nothing until you know what it was yesterday.

4. Make something watch it

Counters nobody reads are decoration. rasdaemon is the current tool; mcelog and edac-utils are both deprecated.

sudo apt install rasdaemon
sudo systemctl enable --now rasdaemon
sudo ras-mc-ctl --error-count
sudo ras-mc-ctl --summary
Label                 CE      UE
mc#0csrow#2channel#0  0       0
mc#0csrow#2channel#1  0       0

On Windows, corrected hardware errors land in the Event Log as WHEA-Logger warnings, and per MemTest86, if ECC is properly enabled in firmware you get them without configuring anything.

Point your existing monitoring at the CE count and alert on any increase. The site’s monitoring comparison covers what to run; any of them can watch a file.

Two ways a working system looks broken

Zeros everywhere on an AMD board. MemTest86 names the culprit: a firmware option called Platform First Error Handling (PFEH). When it is enabled, the platform swallows the errors and the operating system never sees them. Set it to disabled. On boards with a BMC, out-of-band management can intercept reporting the same way, which is worth knowing before you conclude a Supermicro is broken.

A steady drip of correctable errors on healthy RAM. ECC memory has to be written before it can be read, and firmware normally does that at POST. Boards with Quick Boot enabled skip it, and EDAC then reports corrections on memory that was simply never initialised. Disable Quick Boot; boot takes 30 to 60 seconds longer and the phantom errors stop.

What to do when the counter moves

A single correctable error is not an emergency. The distribution says most machines that log one log a handful and carry on.

Escalation is what matters. Google’s data showed a DIMM that has seen a correctable error is 13 to 228 times more likely to see another that same month, and 70 to 80% of uncorrectable errors were preceded by a correctable one in the same or previous month. Facebook flagged a server for repair above 100 correctable errors per week, and its authors calculated that the policy could lower uncorrectable error rate by up to 2.8x compared with waiting for the first uncorrectable event.

Copy the 100-per-week threshold. It is Facebook’s actual repair rule and it is easy to alert on. Google’s paper is clear about what it costs: the absolute probability of an uncorrectable error following a correctable one runs 0.1% to 2.3% per month, so most modules pulled on that rule would have been fine, and its authors said the swap only pays where downtime is expensive enough to cover the false positives. At fleet scale that is an economics question. At home, replacing one $850 stick a decade is not the expensive outcome. Losing the pool is.

And note what the counter does not protect. ECC catches bits. It does not catch you deleting the wrong dataset. The 3-2-1 rule is a separate obligation and ECC does not reduce it by one copy.

Buy the ASRock Rack B650D4U

$299.00 for the board, then a pair of ECC UDIMMs off its published QVL, then verify with EDAC before you put data on it.

Micro-ATX, AM5, four DIMM slots supporting DDR5 ECC/non-ECC UDIMM, and IPMI on board. That last item is why it wins over an ASUS or ASRock consumer board that also lists ECC in its memory line: the same $299 buys documented ECC and out-of-band management, so when the machine stops posting you can read the console instead of carrying a monitor to it. Newegg lists it at $299 today, down from $359; the August 17 feed had it at $323.19. Two readings three weeks apart, and both in that band.

Check NeweggLink checked 2026-09-07Search AmazonLink checked 2026-09-07

On Intel, buy the Supermicro X13SAE-F instead. W680 chipset, 12th through 14th Gen Core, 192GB of unbuffered ECC or non-ECC UDIMM, BMC included. Same logic, different socket. It is the right answer if you already own the CPU.

Do not pay the ECC premium on a consumer board. A board that lists ECC in its memory spec and publishes no ECC QVL is selling you slot compatibility. You can spend $350 a stick and still read None back out of dimm_edac_mode, and the only thing you bought was a slower boot.

Do not buy a used rack server for cheap ECC. Used rack servers are on my regret list — cheap to buy, expensive to run, far louder than the listings suggest, and “quiet” in that market means quiet for a datacenter. Replacing rack gear with small nodes is the best change I have made to the setup. The used enterprise versus mini PC comparison runs that math properly.

And accept the trade if you are buying small. The MS-01 is the mini PC I would buy for a node, and it is the wrong machine for anyone who needs ECC and IPMI. That is not a flaw in the MS-01. It is the thing you give up for four NICs in that volume, and it is worth giving up for a hypervisor node whose state lives somewhere else. It is not worth giving up for the box that holds the array. The MS-01 review covers what it is good at.

ECC is a decision you make once, at motherboard-and-CPU time, and cannot revisit later. Make it deliberately, then spend ninety seconds proving it took.

Frequently asked questions

How do I check if ECC is actually working on Linux?
Read the EDAC subsystem, not the BIOS screen. Run grep -H . /sys/devices/system/edac/mc/mc*/dimm*/dimm_edac_mode. The kernel prints one of a fixed set of strings, and the ones that mean ECC is live are SECDED, S4ECD4ED and their siblings. None means the controller is running without correction even if ECC modules are seated. If /sys/devices/system/edac/mc/ is empty, no EDAC driver bound to your memory controller, so try modprobe amd64_edac, ie31200_edac or igen6_edac for your platform.
Does dmidecode prove ECC is enabled?
No. dmidecode -t 16 reports the Error Correction Type the firmware wrote into the SMBIOS table, and the dmidecode manual says in its own BUGS section that DMI table data is often inaccurate, incomplete or simply wrong. It is a useful first look and it is not evidence. Use it to see whether the firmware claims Single-bit ECC or Multi-bit ECC, then confirm with EDAC, which reads the memory controller itself.
Is DDR5 on-die ECC the same as ECC memory?
No, and retailer titles blur the two. On-die ECC is built into every DDR5 chip to protect the cells inside it. PassMark’s MemTest86 documentation states plainly that it does not provide end-to-end protection, does not detect errors that occur in transmission between the memory controller and the module, and is completely invisible to the CPU. It reports nothing to your operating system, so it can never tell you a DIMM is going bad. Side-band ECC on a real ECC UDIMM is the thing you are buying.
Is ECC worth it for a home NAS or Proxmox box?
Yes if you decide at motherboard-and-CPU time and verify it after boot. It costs a server board rather than a consumer one, roughly $350 more per 32GB module in the August 2026 Newegg feed, and it buys you a counter that goes non-zero before a machine dies rather than during. It is not worth paying for on a board where you cannot read the counters, and it is not a substitute for backups.

Evidence ledger

Last updated
Methodology
This homelab guide was written and edited by Lowell K. Wood IV in St. Louis County, MO. It draws on 10 cited sources, listed below, each checked against the original page on the date above. Full editorial standard: methodology. I did not measure DRAM error rates. Every error-rate figure here comes from two published fleet studies I read in full: Schroeder, Pinheiro and Weber’s SIGMETRICS 2009 paper on Google’s fleet, and Meza, Wu, Kumar and Mutlu’s DSN 2015 paper on Facebook’s. Field names and their meanings come from the DMTF SMBIOS specification, the Linux EDAC sysfs ABI documentation and the EDAC driver source. Reporting behaviour and BIOS gotchas come from PassMark’s MemTest86 ECC documentation. CPU and board support claims come from AMD’s and Supermicro’s product pages and ASRock Rack’s B650D4U user manual; the board’s memory QVL is a separate list on ASRock Rack’s product page and no part number from it is reproduced here. The board price was read live in the Newegg buy box on September 7, 2026; memory prices come from the Newegg affiliate feed deposited August 17, 2026 and are dated.
Update log
  • 2026-09-07 — Last reviewed and updated.
Corrections
Spotted an error or a stale number? Email hello@techfuelhq.com. Confirmed corrections are added to the update log above.

About the author

Written by Lowell K. Wood IV, who builds and runs TechFuelHQ from St. Louis, Missouri.