ECC and DDR5
Advocates for wider adoption of ECC (error-correcting) RAM argue that increasingly dense, failure-prone memory and silent data corruption justify making ECC the default, potentially even through regulation or taxation of non-ECC modules. Others counter that error rates in typical consumer workloads are low, on-die ECC in DDR5 already mitigates many faults, and mandatory ECC would raise costs and restrict choice, especially for gaming and low-end systems. The exchange also touches on technical uncertainties around DDR5’s on-die ECC, the importance of error reporting, market segmentation by CPU and motherboard vendors, and how real-world error anecdotes compare to the lack of large-scale, published reliability data.
Mandating ECC vs. Consumer Choice
- Some argue ECC should be universal, potentially via regulation or taxes on non‑ECC RAM, due to hidden societal costs of silent data corruption.
- Others see this as overreach or “nanny‑statism,” insisting RAM choice should remain individual, except where strong externalities exist (e.g., safety‑critical systems).
- Disagreement over whether unreliable RAM is analogous to defective or hazardous products that justify consumer‑protection laws.
Cost, Availability & Market Segmentation
- ECC UDIMMs are reported as much more expensive and rarer than non‑ECC; RDIMMs can be cheap on secondary markets but need compatible platforms.
- Several comments trace today’s split to deliberate vendor segmentation (server/workstation vs. consumer lines, Xeon vs. mainstream CPUs).
- Some claim that if ECC were mandated/standard, economies of scale and removal of artificial segmentation would shrink the price gap.
Technical Reliability & Error Sources
- Reported ECC error rates vary widely: some see none over years; others see periodic correctable errors and occasional failing DIMMs.
- Causes discussed include cosmic radiation (especially at altitude), rowhammer, bad seating/oxidation, aging modules, electrical noise, and overclocking.
- There is debate over how dominant cosmic rays really are versus design flaws or mechanical/electrical issues.
DDR5 On‑Die ECC vs. End‑to‑End ECC
- Consensus that DDR5 on‑die ECC mainly restores internal cell reliability to older‑generation levels and does not protect the CPU–DIMM link.
- Concern that on‑die ECC is opaque: it corrects silently and typically does not expose error counts to the OS.
- Some worry about complex interactions between on‑die ECC and system‑level ECC; others say DDR5 coding should ensure multi‑bit errors become detectable, but authoritative data is noted as lacking.
Data Integrity, Filesystems, and ZFS
- Multiple comments emphasize that ECC plus checksumming filesystems (e.g., ZFS, Btrfs) provides strong end‑to‑end protection and aids diagnosis of failing hardware.
- Others argue non‑ECC is often acceptable because many software layers already use checksums, retries, and redundancy; they view catastrophic ECC‑less corruption as rare.
- One long‑running myth that “ECC is mandatory for ZFS” is explicitly challenged and linked to counter‑documentation.