Qwen3.8-Max: A New Bar for Coding and Cowork

Alibaba’s release of Qwen3.8-Max — a high-end coding- and “cowork”-focused AI model with open weights promised, plus an upcoming 27B variant — is being seen as a major step forward for open-weight, China-led models that rival Western closed systems in capability and price. Commenters compare Qwen to Kimi K3, DeepSeek, GLM and frontier models like Claude and GPT, debating where it now sits for coding, local deployment, and long-horizon agentic work, and how cheap, powerful Chinese models pressure US labs on both pricing and regulation. A recurring theme is whether local models on consumer or prosumer hardware can be economically justified versus cloud APIs, and what widespread automation means for software jobs, enterprise lock‑in, and the future AI “moat” (compute, data, or tooling).

Model release & openness

  • Qwen3.8-Max is described as the new flagship, with open weights for the Max-class model promised “next week”; a 27B open‑weight variant is also announced and widely anticipated.
  • Users clarify that July’s “Max-Preview” was an earlier checkpoint; this is the full release.
  • Many hope the license matches earlier Apache-style Qwen releases; some contrast this with Kimi K3’s more restrictive commercial terms and GLM 5.2’s MIT license.

Performance, pricing & comparison

  • Qwen3.6‑27B and 35B-A3B are already popular local models, often described as the best “daily drivers” for coding and agents; 27B dense is frequently called smarter and more coherent than the 35B MoE, which is faster and better for agents.
  • Qwen3.8-Max is reported to perform near top closed models for coding and visual tasks in early anecdotes, though harness/tooling issues (timeouts, flaky desktop apps) are noted.
  • Pricing on Qwen Cloud is cited as $2/M input, $6/M output, $0.25/M cached tokens—competitive versus Kimi K3 and DeepSeek; some say DeepSeek Flash remains cheaper for many workloads.

Local vs cloud, hardware & economics

  • Many run Qwen 3.6 (27B/35B) locally on Macs, RTX GPUs, Strix Halo, or mixed CPU+GPU setups; quantization (4–8 bit) and MoE are key to fitting big models into 16–64GB RAM.
  • Some argue local models are mainly for privacy, control and hobby use; pure cost per token often favors cheap hosted Chinese models, especially with caching.
  • Others claim that for enterprises with heavy usage, self-hosting strong open‑weights (Chinese or otherwise) can undercut frontier API costs and provide leverage in negotiations.

Coding, agents & workflows

  • Qwen models are heavily used for agentic coding with harnesses (e.g., Pi-based systems); long‑horizon autonomous runs like “oh‑my-cli/oh‑my-pi” impress some but also highlight reliability issues.
  • Several report that AI increased workload: more tasks, done faster, with expectations ratcheting up rather than time freed for hobbies.
  • There’s debate on whether specialized “coding-only” models would be meaningfully smaller; most argue general world knowledge measurably improves coding ability.

Geopolitics, regulation & moats

  • Many see Chinese open‑weight models (Qwen, DeepSeek, Kimi, GLM) as driving a race to the bottom on inference prices and eroding the moat of US closed labs.
  • Others argue moats now lie in compute access, proprietary RLHF data, enterprise integration, regulatory capture, and harness‑level lock‑in (non‑portable session state).
  • Speculation about US or EU moves to restrict Chinese models or open weights is common; feasibility and enforceability are hotly debated.