Apple wants AI to run directly on its hardware instead of in the cloud

Apple’s push to run generative AI directly on iPhones and other devices is seen as a way to improve privacy, reduce reliance on cloud infrastructure, and enable faster, more reliable features like a much better Siri or on-device image and text processing. Commenters are split on whether current hardware—especially limited RAM—is sufficient for useful large language models, but many note that smaller models already run decently on recent iPhones and iPads. Beyond technical hurdles, people see this strategy as both a cost-saving move for Apple and a potential driver of future hardware upgrade cycles in an already saturated smartphone market.

On-device vs Cloud AI

  • Many see on-device inference as the “next step” for phones: lower latency, offline use, and better privacy.
  • Others note that cloud models will remain more capable due to larger model sizes and data access; local will often be “good enough,” not best.
  • Hybrid visions are common: base model and personal context on-device, with optional cloud calls for current events or heavy tasks.

Siri and Voice Assistants

  • Strong demand for a significantly improved Siri; some say a working, LLM-grade Siri alone would justify a phone upgrade.
  • Frustration is widespread: inconsistent behavior, broken “incantations,” random regressions, and poor multi-user handling.
  • Some argue an LLM could fix robustness and understanding; others want Apple to first fix basic stability before adding complexity.

Hardware, Models, and Performance

  • Debate over Apple’s historically low RAM: 8 GB is seen as tight for modern LLMs, though current high-end iPads/Macs can already run 7B–13B models locally.
  • Bandwidth/RAM, not storage, is seen as the main bottleneck; models need to fit in memory.
  • Apple’s recent research on low‑memory LLM inference is cited as an attempt to work around these constraints.
  • Some expect dedicated AI co-processors and possibly an AI-focused device tier.

Privacy, Business Model, and Costs

  • On-device AI aligns with Apple’s privacy messaging and reduces their cloud compute bills.
  • Several note that Apple prefers selling high‑margin hardware over running costly inference for free.
  • There’s interest in optional cloud-based enhancements, but questions about who would pay and how (subscriptions, bundles).

Ecosystem, Openness, and Regulation

  • There’s skepticism that Apple will allow arbitrary system-level AI models; expectation is a single “official” assistant.
  • Counterpoint: current APIs (Core ML, Neural Engine) already allow third-party on-device models, though not as Siri replacements.
  • EU pressure on app store openness is discussed; how far this will extend to AI models and assistants is unclear.

Existing On-device AI & UX

  • Users report strong experiences with on-device features: photo content search, Live Text, image description, and on-watch Siri processing.
  • These “invisible” AI features are cited as proof Apple has been investing in local ML for years, even if not branded as generative AI.