Key Takeaways
Gemini vs. ChatGPT vs. Claude vs. Grok vs. Perplexity! (The Best Way To Use Each One)

- GPT-5.6 Sol produced the most accurate reproductions of both the Mona Lisa and Starry Night targets
- Claude Fable 5 took significantly longer and cost more but delivered worse visual output than GPT-5.6
- Grok 4.5 struggled with basic drawing tasks, and open-weight models returned blank canvases
Give four frontier AI models a blank canvas, a set of colored-pencil tools, and the Mona Lisa as a target. Who wins? TryAI built exactly this experiment, running GPT-5.6 Sol, Claude Fable 5, Grok 4.5, and Gemini 3.6 Flash through 28 drawings. The results expose real gaps between frontier models that benchmarks often hide.
How the drawing arena works
The harness, open-sourced at github.com/hershalb/canvas-arena, gives every model identical tools: set_color, set_brush, set_pressure, draw, smudge, erase, and a view_canvas function to see their own work. No solid fills exist. Tone and color come from layering strokes, like an actual pencil. Models could call view_target to reference the original image at any point.
For the two reproduction tasks, the Mona Lisa and Van Gogh's Starry Night, each output was scored using SSIM (structural similarity index). The five open-ended prompts had no target and no score. The prompts ranged from "an elderly fisherman's weathered face" to "a cozy cabin interior with a lit fireplace."
Mona Lisa: the first real test

The target image was straightforward. Reproduce da Vinci's masterpiece using only colored-pencil strokes, blending, and erasing. Here's how each model performed.

GPT-5.6 Sol captured the composition cleanly. The face proportions hold, the sfumato effect around the eyes translates reasonably well, and the background landscape reads correctly. This was the highest-scoring reproduction.

Claude Fable 5 took a different approach. It spent far more time iterating, calling view_canvas repeatedly, planning extensively. The result? Worse output, more tokens burned, higher cost. The face structure is less accurate, the tonal range narrower.

Grok 4.5 struggled. The proportions are off, the layering technique inconsistent. For a task that should be "basic" for frontier models, this was rough. The TryAI team noted that several open-weight models they tested returned blank canvases entirely.

Gemini 3.6 Flash landed between Claude and Grok. Competent but not exceptional. The color palette was accurate, but fine detail work lagged behind GPT-5.6.
Starry Night: testing complex texture

Van Gogh's swirling brushwork presents a different challenge. The patterns are more abstract, the color transitions more dramatic.

GPT-5.6 Sol again led the field. The swirl patterns are recognizable, the village silhouette intact, the color gradients convincing. Not a perfect reproduction, but clearly the strongest attempt.

Claude Fable 5 improved relative to its Mona Lisa attempt but still trailed GPT-5.6. The swirls are there, the composition roughly correct, but the execution lacks the precision of the OpenAI model.
What this reveals about model costs
The TryAI team highlighted a crucial finding: Claude Fable 5 "almost always took much longer than the others, for far more money, and here produced worse output." This matters for founders building AI-powered products. Token cost and inference time translate directly to user experience and unit economics.
Claude Fable 5 has beaten GPT-5.6 on other tasks, including TryAI's previous music-video challenge. But on this drawing benchmark, the cost-performance ratio favored OpenAI.
| Model | Mona Lisa Performance | Cost/Time | Strengths |
|---|---|---|---|
| GPT-5.6 Sol | Highest SSIM | Moderate | Composition, detail, efficiency |
| Claude Fable 5 | Lower SSIM | Highest | Iterative refinement, planning |
| Grok 4.5 | Lowest SSIM | Low | None observed in this task |
| Gemini 3.6 Flash | Mid-range SSIM | Low | Color accuracy, speed |
If you're evaluating API options for AI-powered features
Why open-ended drawing tasks matter
TryAI addressed criticism from their previous music-video benchmark. These tests aren't meant to prove models are "creative" or "artistic." They serve a different purpose.
First, they cut through benchmaxxing. Many open-weight models post impressive numbers on standard evals but fail at open-ended tasks. Several returned blank canvases when asked to draw anything. That gap between benchmark scores and real capability matters when you're shipping a product.
Second, they reveal the frontier-versus-open gap clearly. Cheaper open-weight models can replace frontier models for execution work. But for tasks requiring planning, iteration, and self-correction, the gap remains wide.
What about Kimi K3?
TryAI noted they plan to test Kimi K3 once it's fully open-sourced. The implication: open-weight models may close this gap soon. For founders watching AI infrastructure costs, that timeline matters.
Related context on how AI changes software economics
The prompt-only drawings
Five additional prompts tested each model without a target image: an elderly fisherman's face, a sunset over calm ocean, a red rose with Rembrandt lighting, a sleeping tabby cat, and a cabin interior with fireplace. Without SSIM scoring, these become subjective. But the patterns held. GPT-5.6 Sol produced the most coherent compositions. Claude Fable 5 over-iterated. Grok 4.5 remained inconsistent.
Logicity's Take
For founders building products with AI-generated visuals, this test reveals something important: frontier model rankings shift by task. Claude beats GPT on some workloads, loses badly on others. The cost difference isn't marginal. If you're prototyping with one model's API, test your actual use case before committing. Tools like [Zapier](https://logicity.in/r/zapier) or [Make](https://logicity.in/r/make) can help you A/B test different model endpoints without rebuilding your entire pipeline.
Disclosure
Some links in this post are affiliate links — Logicity earns a commission if you sign up, at no extra cost to you. We only link products we have used or actively recommend.
Frequently Asked Questions
Which AI model is best for drawing tasks?
In TryAI's colored-pencil arena, GPT-5.6 Sol produced the highest-quality reproductions of both the Mona Lisa and Starry Night, with lower costs than Claude Fable 5.
Can open-weight AI models compete with frontier models on creative tasks?
Not yet. Several open-weight models TryAI tested returned blank canvases. The gap between benchmark scores and real-world creative capability remains significant.
How much does Claude Fable 5 cost compared to GPT-5.6?
TryAI reported Claude Fable 5 took much longer and cost significantly more than GPT-5.6 Sol while producing worse output on the drawing tasks.
Is the canvas-arena drawing test open source?
Yes. The harness is available at github.com/hershalb/canvas-arena. You can point it at any target image or text prompt.
Need Help Implementing This?
Evaluating AI models for your product? Logicity's consulting team helps founders benchmark model performance against their actual use cases. Reach out at consulting@logicity.in.
Source: Hacker News: Best / TryAI
Manaal Khan
Tech & Innovation Writer
Produced with AI assistance and reviewed by the Logicity editorial team. Learn more in our Editorial Policy.
Related Articles
More in Startups & Innovation
Redwood Materials Layoffs Signal Battery Industry Pivot
Redwood Materials cut 135 jobs (10% of staff) while pivoting toward energy storage, just three months after raising $425M at a $6B+ valuation. For executives watching the battery and EV supply chain, this restructuring reveals where smart money is heading as automotive electrification plans cool down.





