All posts

GPT-5.6 vs Claude vs Grok vs Gemini: which draws best?

Manaal KhanJuly 23, 2026 at 8:16 PM6 min read
GPT-5.6 vs Claude vs Grok vs Gemini: which draws best?

Key Takeaways

Gemini vs. ChatGPT vs. Claude vs. Grok vs. Perplexity! (The Best Way To Use Each One)

GPT-5.6 vs Claude vs Grok vs Gemini: which draws best?
Source: Hacker News: Best
  • GPT-5.6 Sol produced the most accurate reproductions of both the Mona Lisa and Starry Night targets
  • Claude Fable 5 took significantly longer and cost more but delivered worse visual output than GPT-5.6
  • Grok 4.5 struggled with basic drawing tasks, and open-weight models returned blank canvases

Give four frontier AI models a blank canvas, a set of colored-pencil tools, and the Mona Lisa as a target. Who wins? TryAI built exactly this experiment, running GPT-5.6 Sol, Claude Fable 5, Grok 4.5, and Gemini 3.6 Flash through 28 drawings. The results expose real gaps between frontier models that benchmarks often hide.

Advertisements

How the drawing arena works

The harness, open-sourced at github.com/hershalb/canvas-arena, gives every model identical tools: set_color, set_brush, set_pressure, draw, smudge, erase, and a view_canvas function to see their own work. No solid fills exist. Tone and color come from layering strokes, like an actual pencil. Models could call view_target to reference the original image at any point.

For the two reproduction tasks, the Mona Lisa and Van Gogh's Starry Night, each output was scored using SSIM (structural similarity index). The five open-ended prompts had no target and no score. The prompts ranged from "an elderly fisherman's weathered face" to "a cozy cabin interior with a lit fireplace."

Mona Lisa: the first real test

Mona Lisa target
Mona Lisa target

The target image was straightforward. Reproduce da Vinci's masterpiece using only colored-pencil strokes, blending, and erasing. Here's how each model performed.

Mona Lisa by GPT-5.6 Sol
Mona Lisa by GPT-5.6 Sol

GPT-5.6 Sol captured the composition cleanly. The face proportions hold, the sfumato effect around the eyes translates reasonably well, and the background landscape reads correctly. This was the highest-scoring reproduction.

Mona Lisa by Claude Fable 5
Mona Lisa by Claude Fable 5

Claude Fable 5 took a different approach. It spent far more time iterating, calling view_canvas repeatedly, planning extensively. The result? Worse output, more tokens burned, higher cost. The face structure is less accurate, the tonal range narrower.

Mona Lisa by Grok 4.5
Mona Lisa by Grok 4.5

Grok 4.5 struggled. The proportions are off, the layering technique inconsistent. For a task that should be "basic" for frontier models, this was rough. The TryAI team noted that several open-weight models they tested returned blank canvases entirely.

Mona Lisa by Gemini 3.6 Flash
Mona Lisa by Gemini 3.6 Flash

Gemini 3.6 Flash landed between Claude and Grok. Competent but not exceptional. The color palette was accurate, but fine detail work lagged behind GPT-5.6.

Starry Night: testing complex texture

Starry Night target
Starry Night target

Van Gogh's swirling brushwork presents a different challenge. The patterns are more abstract, the color transitions more dramatic.

Starry Night by GPT-5.6 Sol
Starry Night by GPT-5.6 Sol

GPT-5.6 Sol again led the field. The swirl patterns are recognizable, the village silhouette intact, the color gradients convincing. Not a perfect reproduction, but clearly the strongest attempt.

Starry Night by Claude Fable 5
Starry Night by Claude Fable 5

Claude Fable 5 improved relative to its Mona Lisa attempt but still trailed GPT-5.6. The swirls are there, the composition roughly correct, but the execution lacks the precision of the OpenAI model.

What this reveals about model costs

The TryAI team highlighted a crucial finding: Claude Fable 5 "almost always took much longer than the others, for far more money, and here produced worse output." This matters for founders building AI-powered products. Token cost and inference time translate directly to user experience and unit economics.

Claude Fable 5 has beaten GPT-5.6 on other tasks, including TryAI's previous music-video challenge. But on this drawing benchmark, the cost-performance ratio favored OpenAI.

ModelMona Lisa PerformanceCost/TimeStrengths
GPT-5.6 SolHighest SSIMModerateComposition, detail, efficiency
Claude Fable 5Lower SSIMHighestIterative refinement, planning
Grok 4.5Lowest SSIMLowNone observed in this task
Gemini 3.6 FlashMid-range SSIMLowColor accuracy, speed
Also Read
6 OpenAI-compatible APIs compared: which drop-in fits?

If you're evaluating API options for AI-powered features

Advertisements

Why open-ended drawing tasks matter

TryAI addressed criticism from their previous music-video benchmark. These tests aren't meant to prove models are "creative" or "artistic." They serve a different purpose.

First, they cut through benchmaxxing. Many open-weight models post impressive numbers on standard evals but fail at open-ended tasks. Several returned blank canvases when asked to draw anything. That gap between benchmark scores and real capability matters when you're shipping a product.

Second, they reveal the frontier-versus-open gap clearly. Cheaper open-weight models can replace frontier models for execution work. But for tasks requiring planning, iteration, and self-correction, the gap remains wide.

What about Kimi K3?

TryAI noted they plan to test Kimi K3 once it's fully open-sourced. The implication: open-weight models may close this gap soon. For founders watching AI infrastructure costs, that timeline matters.

Also Read
Services-as-software: 5 ways SaaS firms should adapt

Related context on how AI changes software economics

The prompt-only drawings

Five additional prompts tested each model without a target image: an elderly fisherman's face, a sunset over calm ocean, a red rose with Rembrandt lighting, a sleeping tabby cat, and a cabin interior with fireplace. Without SSIM scoring, these become subjective. But the patterns held. GPT-5.6 Sol produced the most coherent compositions. Claude Fable 5 over-iterated. Grok 4.5 remained inconsistent.

ℹ️

Logicity's Take

For founders building products with AI-generated visuals, this test reveals something important: frontier model rankings shift by task. Claude beats GPT on some workloads, loses badly on others. The cost difference isn't marginal. If you're prototyping with one model's API, test your actual use case before committing. Tools like [Zapier](https://logicity.in/r/zapier) or [Make](https://logicity.in/r/make) can help you A/B test different model endpoints without rebuilding your entire pipeline.

ℹ️

Disclosure

Some links in this post are affiliate links — Logicity earns a commission if you sign up, at no extra cost to you. We only link products we have used or actively recommend.

Frequently Asked Questions

Which AI model is best for drawing tasks?

In TryAI's colored-pencil arena, GPT-5.6 Sol produced the highest-quality reproductions of both the Mona Lisa and Starry Night, with lower costs than Claude Fable 5.

Can open-weight AI models compete with frontier models on creative tasks?

Not yet. Several open-weight models TryAI tested returned blank canvases. The gap between benchmark scores and real-world creative capability remains significant.

How much does Claude Fable 5 cost compared to GPT-5.6?

TryAI reported Claude Fable 5 took much longer and cost significantly more than GPT-5.6 Sol while producing worse output on the drawing tasks.

Is the canvas-arena drawing test open source?

Yes. The harness is available at github.com/hershalb/canvas-arena. You can point it at any target image or text prompt.

ℹ️

Need Help Implementing This?

Evaluating AI models for your product? Logicity's consulting team helps founders benchmark model performance against their actual use cases. Reach out at consulting@logicity.in.

Source: Hacker News: Best / TryAI

M

Manaal Khan

Tech & Innovation Writer

Produced with AI assistance and reviewed by the Logicity editorial team. Learn more in our Editorial Policy.