DeepSeek V4 Pro 0813
DeepSeek’s new V4 Pro 0813 model is being compared against both its own V4 Flash variant and US “frontier” models like Claude Opus, Fable, and OpenAI’s Sol/Luna: benchmarks put it just below the very top systems but at a tiny fraction of their per‑task cost, especially when DeepSeek’s aggressive prompt caching is used. Many engineers report that Flash still offers the best price‑performance for day‑to‑day coding, while Pro can be worth it for planning, auditing, and harder reasoning tasks, though results vary heavily with the chosen agent harness and workflow. Commenters also weigh trade‑offs around privacy, Chinese model adoption in Western enterprises, and DeepSeek’s announced move to higher, peak/off‑peak pricing, which may still leave it far cheaper than Western competitors.
Overall positioning vs other models
- Benchmarks show DeepSeek V4 Pro 0813 near Fable 5 / Opus 5 class and above GLM 5.2, with Flash close behind.
- Users compare it to Qwen3.8-max, Kimi-K3, GPT-5.6 (Luna/Sol), Grok 4.6, MiMo, GLM 5.2: general consensus is “frontier-adjacent” but not clearly best-in-class.
- Some report Pro as weaker than Luna or Kimi on their tasks; others say it beats Luna once caching and cost are included.
Pricing, caching, and subscriptions
- Raw token prices are very low; with heavy cache reads, some claim 10–60x cheaper per task than Opus, and cheaper than Luna for coding harnesses.
- DeepSeek is introducing peak/off‑peak pricing and warning of “significant” future increases; confusion remains about timing and exact new rates.
- Debate: API token buyers vs heavily subsidized ChatGPT/Claude subscriptions. Many note that subscriptions can effectively undercut all per-token pricing.
Pro vs Flash and real‑world coding
- Benchmarks: Pro ~5 points above Flash on coding/agent tasks, but many find Flash “good enough” and far cheaper.
- Some workflows: “Pro plans, Flash implements” or the reverse; others use frontier models (Sol/Fable/Opus) to plan/review and DeepSeek to implement.
- Several note Pro feels slower and not always worth it over Flash; others say Pro is more reliable per task and less verbose.
- Strong consensus that harness/tooling matters a lot: results differ drastically across Pi, OpenCode, Codex, etc.
Quality, hallucinations, and domains
- Mixed reports: some call Pro “unreliable” at pass@1 but strong at pass@3; others say it matches or beats Sol/Fable in their multi‑shot agent loops.
- ArtificialAnalysis hallucination metrics are reportedly bad for DeepSeek models; some users confirm more hallucinations in non‑coding or multilingual prose.
- For security review, performance tuning, research/evaluation, and agentic loops, several users are impressed by Pro’s capability per dollar.
Privacy, politics, and hosting
- Concerns about DeepSeek’s data retention and training on prompts; some refuse to benchmark until a zero‑retention provider exists.
- Others dismiss these worries, or even prefer data outside US jurisdiction.
- Enterprises may avoid Chinese-origin models due to regulatory/political risk, even if hosted in the US/EU or run locally.
Ecosystem and deployment
- Weights for Pro 0813 are not yet widely available; Flash can already be run locally and is praised as “too cheap to meter.”
- Users discuss alternatives to OpenRouter (other routers, direct API use) for better caching, privacy controls, or lower fees.