ECC and DDR5

Advocates for wider adoption of ECC (error-correcting) RAM argue that increasingly dense, failure-prone memory and silent data corruption justify making ECC the default, potentially even through regulation or taxation of non-ECC modules. Others counter that error rates in typical consumer workloads are low, on-die ECC in DDR5 already mitigates many faults, and mandatory ECC would raise costs and restrict choice, especially for gaming and low-end systems. The exchange also touches on technical uncertainties around DDR5’s on-die ECC, the importance of error reporting, market segmentation by CPU and motherboard vendors, and how real-world error anecdotes compare to the lack of large-scale, published reliability data.

Mandating ECC vs. Consumer Choice

  • Some argue ECC should be universal, potentially via regulation or taxes on non‑ECC RAM, due to hidden societal costs of silent data corruption.
  • Others see this as overreach or “nanny‑statism,” insisting RAM choice should remain individual, except where strong externalities exist (e.g., safety‑critical systems).
  • Disagreement over whether unreliable RAM is analogous to defective or hazardous products that justify consumer‑protection laws.

Cost, Availability & Market Segmentation

  • ECC UDIMMs are reported as much more expensive and rarer than non‑ECC; RDIMMs can be cheap on secondary markets but need compatible platforms.
  • Several comments trace today’s split to deliberate vendor segmentation (server/workstation vs. consumer lines, Xeon vs. mainstream CPUs).
  • Some claim that if ECC were mandated/standard, economies of scale and removal of artificial segmentation would shrink the price gap.

Technical Reliability & Error Sources

  • Reported ECC error rates vary widely: some see none over years; others see periodic correctable errors and occasional failing DIMMs.
  • Causes discussed include cosmic radiation (especially at altitude), rowhammer, bad seating/oxidation, aging modules, electrical noise, and overclocking.
  • There is debate over how dominant cosmic rays really are versus design flaws or mechanical/electrical issues.

DDR5 On‑Die ECC vs. End‑to‑End ECC

  • Consensus that DDR5 on‑die ECC mainly restores internal cell reliability to older‑generation levels and does not protect the CPU–DIMM link.
  • Concern that on‑die ECC is opaque: it corrects silently and typically does not expose error counts to the OS.
  • Some worry about complex interactions between on‑die ECC and system‑level ECC; others say DDR5 coding should ensure multi‑bit errors become detectable, but authoritative data is noted as lacking.

Data Integrity, Filesystems, and ZFS

  • Multiple comments emphasize that ECC plus checksumming filesystems (e.g., ZFS, Btrfs) provides strong end‑to‑end protection and aids diagnosis of failing hardware.
  • Others argue non‑ECC is often acceptable because many software layers already use checksums, retries, and redundancy; they view catastrophic ECC‑less corruption as rare.
  • One long‑running myth that “ECC is mandatory for ZFS” is explicitly challenged and linked to counter‑documentation.