Our next-generation model: Gemini 1.5
Google’s announcement of its Gemini 1.5 Pro AI model, featuring a Mixture-of-Experts architecture and a context window tested up to 10 million tokens, is seen as a potential leap for tasks like codebase analysis, document understanding, and multimodal (video/audio/text) reasoning. Commenters are intrigued by the long-context capabilities and possible impacts on techniques like retrieval-augmented generation, but heavily criticize Google’s confusing product naming, region and waitlist restrictions, and prior overhyped or constrained releases. Many conclude that real judgment will have to wait until broad, hands-on access is available and pricing, speed, and reliability are clearer.
Branding, Versions, and Product Tiers
- Many find the naming matrix confusing: Gemini 1.0 vs 1.5, and size tiers Nano / Pro / Ultra overlapped with consumer products “Gemini” (free) and “Gemini Advanced” (paid).
- Rough consensus:
- Models: Nano (on-device), Pro (mid), Ultra (largest).
- Versions: 1.0, 1.5, etc.
- Current mapping: Free chat ≈ Gemini 1.0 Pro; Gemini Advanced ≈ Gemini 1.0 Ultra; this announcement is Gemini 1.5 Pro only.
- People criticize Google for unclear branding compared with simpler versioning elsewhere.
Long Context Window & RAG
- Headline feature: up to 1M tokens in production, 10M in research, across text, audio, image, and video.
- Demos and tech report claim near-perfect “needle in a haystack” retrieval up to millions of tokens and impressive tasks like querying a 44‑minute film or 1400‑page book, and learning a low-resource language (Kalamang) from its grammar book.
- Some argue this could drastically simplify or replace many Retrieval-Augmented Generation (RAG) stacks; others say RAG will still matter for speed, cost, accuracy, and very large corpora.
Architecture, Performance, and Benchmarks
- 1.5 Pro uses a new Mixture-of-Experts architecture and is claimed to match or surpass Gemini 1.0 Ultra on many benchmarks.
- Discussion notes that published comparisons with GPT‑4 are selective; some infer 1.5 Pro ≈ GPT‑4, others are skeptical and want independent tests.
- The report highlights data contamination problems on coding benchmarks like HumanEval; they propose an alternative benchmark (Natural2Code).
Compute, Latency, and Cost
- Several commenters worry about the economics: sending 1M–10M tokens per request could be very expensive and slow (demo answers reportedly ~60 seconds for large contexts).
- Even if technically possible, many expect RAG or other filtering to remain important for price/performance.
Access, Rollout, and Region Locking
- 1.5 Pro is only in limited preview via AI Studio / Vertex AI, often behind waitlists and allowlists.
- EU/UK users in particular report they can’t access AI Studio or Ultra APIs despite paid subscriptions.
- Many express frustration with “announce now, waitlist later,” saying it destroys enthusiasm and trust.
Safety, Guardrails, and Usefulness
- Multiple reports that current Gemini chat is over-restrictive (moral lectures, refusals on benign coding and everyday tasks, extra limits for minors), making it feel less useful than ChatGPT.
- Tension noted between Google’s emphasis on “safety at the core” and developer desire for less-hobbled models.
Trust in Google and Strategic Position
- Some are impressed by the long-context and multimodal demos; others distrust Google after earlier heavily edited Gemini videos and years of “beta/waitlist” products.
- Debate over whether this shows Google has caught up with or surpassed OpenAI, or whether OpenAI still leads on practical, widely available capabilities.