Training for one trillion parameter model backed by Intel and US govt has begun

A US government–backed project with Intel has begun training a trillion-parameter AI model on the Aurora supercomputer, aiming to rival or approach systems like GPT‑4 and advance scientific research. Commenters debate whether ever-larger models are the right path versus new architectures and training methods, and question if there is enough high‑quality data to scale much further. The effort also raises concerns about state control of powerful AI, potential use of surveilled or classified data, openness of the resulting model, and broader societal risks versus benefits of continued AI escalation.

Model scale, architecture, and comparisons

  • Some note 1T parameters may only put this in GPT‑4’s class, not beyond it; GPT‑4’s true size and whether it’s Mixture‑of‑Experts (MoE) remain unconfirmed and debated.
  • Several point out that all currently known 1T+ models are MoE; whether this new one is dense or MoE is unclear but crucial for comparisons.
  • Scaling hypothesis is widely referenced: so far, “bigger model + more data + more compute” reliably improves capability, though many expect eventual limits.
  • Others argue future gains will require architectural changes and better data efficiency, not just more parameters.

Data, training, and human comparison

  • Debate on whether we’re “running out of data”; some claim major labs can still find or generate plenty, others see limits.
  • Back‑of‑the‑envelope estimates compare human sensory input (hundreds of TB–PB scale) to LLM training corpora, suggesting humans may experience vastly more raw input but in more interactive and multimodal ways.
  • Several stress that humans learn via active interaction and feedback with the real world, not just text; some suggest future AI must similarly act in environments to reach higher intelligence.

Supercomputers and hardware

  • Intel/Argonne’s Aurora system is seen as a showcase for Intel GPUs and massive parallelism.
  • There’s discussion over whether traditional supercomputer architectures (e.g., HPC interconnects, FP64 bias) are ideal for LLM training, versus large GPU clusters; opinions are mixed.

Benefits and scientific goals

  • Many are optimistic about a domain‑specific, science‑trained model: better literature search, autonomous discovery, and acceleration of research.
  • Some see this as one of the most compelling public‑sector AI projects.

Risks, ethics, and governance

  • Strong “AI doomer” perspectives argue AI is net harmful, likening it to chemical weapons and broader technological harms (climate, plastics, industrial agriculture).
  • Others counter that technology (including AI) brings large benefits and that problems typically stem from misuse, economics, and governance, not science itself.
  • Significant debate over whether governments or corporations are the lesser evil to control powerful AI, with extensive historical arguments about state vs corporate atrocities.
  • Concerns raised about surveillance data: some speculate intelligence agencies could train powerful, private models on decades of communications; this is noted as alarming but unverified.

Openness, access, and security

  • Questions about whether model weights could be released or FOIA’d; most expect “national security” classification.
  • Tension between desires for open models and worries about handing advanced capabilities to adversaries.