Qwen 3.8 Omni Flash
Qwen 3.8 Omni Flash is drawing attention as a fast, low-cost alternative to models like Gemini Omni, with some users reporting strong performance—especially from variants like Flash Next and Max—despite heavy hardware requirements for the largest versions. Commenters highlight trade-offs between speed, reasoning quality, aggressive RL-tuning, and quantization or serving bugs that can cause “weird” behavior. The thread also broadens into concerns about model pricing, Europe’s relative weakness in building foundation models compared to China and the US, and the growing difficulty of choosing among an expanding ecosystem of LLMs.
Qwen 3.8 models & open weights
- Users appreciate Qwen’s range of model sizes, especially very small ones for experimentation.
- Some perceive a slowdown in open-weight releases (e.g., uncertainty about Qwen4, limited sizes like 27B and 125B), though others note that major recent models like 3.8 Max and 3.8 Flash Next have been released as weights.
Performance, usability & hardware constraints
- 125B / Flash Next is praised as “very strong” in quality, even at 3-bit quantization, but considered “mostly unusable” for many due to high RAM requirements (~128GB) and slow throughput.
- 3.8 Max is seen as very grounded, reliable, and “boring in a good way,” but slow and only easily accessible via Alibaba, whose pricing is described as stingy.
- Some users prefer Flash Next over 27B and have dropped the latter entirely on their hardware.
Weird reasoning & serving/quantization issues
- Multiple people report Flash Next producing odd “meta” thoughts (e.g., claiming the user didn’t ask anything, or random philosophical musings) between tool calls.
- These issues are suspected to be due to serving bugs or bad quantizations, though others see similar behavior even with seemingly good setups.
- Some report long, random stalls or excessively long chains-of-thought that resolve after retrying.
RL and “over-tuning” concerns
- There is worry that aggressive reinforcement learning for agentic coding could make models overly eager, doing things without clarifying user intent.
- Some see newer frontier models as examples where optimization for benchmarks harms everyday usefulness.
Cost vs Gemini and other models
- Qwen 3.8 Omni Flash is claimed (by the post) to approach Gemini’s multimodal quality at ~10× lower per-token cost.
- Commenters note true cost depends heavily on tokens used per task, cache behavior, and task type; in some benchmarks, far more expensive models can be cheaper per completed task.
Model choice & tool support
- Users feel overwhelmed by model choice; suggestions include: start with a cheap, good-enough model; avoid chasing every new release; and use tools like OpenRouter’s “auto model” or comparison sites, with caveats for production pipelines.
China vs Europe in AI
- Debate over why Chinese labs ship strong models while Europe lags: cited factors include government direction, capital intensity, risk appetite, regulation, and education.
- Others counter that Europe has notable labs, companies, and chip-related infrastructure, but less “hustle” and more regulation and risk aversion.
Branding, docs & UX
- Some criticize confusing product names (Flash/Omni/Pro/Ultra) and prefer clearer versioning.
- The official blog is described as technically heavy, sluggish, and full of flashy but uninformative media, with poor readability.