"Drawing" the Mona Lisa with GPT-5.6, Claude, Gemini, and Grok

LLMs like GPT-5.6, Claude, Gemini, and Grok were tasked with “drawing” images such as the Mona Lisa and Starry Night using only simple digital painting tools, revealing wildly different levels of control, style, and cost-efficiency. Commenters note that despite looking childish or symbolic—much like a beginner’s art—the results show surprising emergent capabilities in tool use and visual reasoning, with GPT-5.6 often outperforming far more expensive setups. The thread also raises questions about evaluation methods, prompt design, model specialization, and whether human-like analogies (e.g., comparing models to children or parrots) meaningfully describe current AI progress.

Purpose and validity of the test

  • Several commenters see the piece as primarily a marketing/engagement article rather than a rigorous comparison.
  • Critiques include: identical prompts across very different models, no tailoring to each model’s “style,” no repeated runs for consistency, and no “undo/revert” tool, which biases against backtracking strategies.
  • Some argue that in real deployments, you’d adapt the harness and prompting to each model, so the head‑to‑head comparison is inherently limited.

Are LLMs suitable for drawing?

  • Multiple people stress these are language models using a drawing tool, not native image generators, so calling the results “useless” is seen as unfair.
  • Others still find the output unimpressive and note that specialized image models or image-capable chat systems perform much better.
  • There’s interest in better harnesses: more tools, zoom/cropping, palette management, “undo,” and using better similarity metrics (e.g., vision-model embeddings instead of SSIM/RMSE).

Model behavior and style

  • Many agree GPT-5.6 “Sol” is the clear winner: more coherent, efficient, and “charming,” especially on the rose and Starry Night.
  • Grok’s drawings are widely described as disturbingly bad, surreal, or like the infamous botched “restoration” of a Jesus painting; some still find them darkly funny or emotionally evocative.
  • Claude and other models are seen as middling, sometimes over-focusing on blending or misusing the limited toolset.

“Childlike” drawing and anthropomorphism

  • A recurring theme is that outputs look like work from beginner or child artists: “symbol drawing” based on icons (“blue = glass”) rather than light, form, and value.
  • Some see this as eerily human-like and suggest it parallels human artistic development; others push back, saying models don’t actually “develop,” labs just train better versions.
  • There’s debate over anthropomorphizing LLMs vs treating them as mere tools, with some frustration at overused “stochastic parrot” analogies.

Costs, efficiency, and economics

  • Commenters highlight the large cost gap, especially between GPT-5.6 Sol and Fable, noting Sol achieves better results with far fewer tokens and dollars.
  • Some report LLM usage feeling cheaper and better over time; others say they still pay more overall, attributing that to increased usage rather than per‑unit cost.