English

Pricing & LimitsGPT-5.6 SolClaude Fable 5Grok 4.5Gemini 3.6 FlashHarsalb

Four Models Including GPT-5.6 and Claude Fable 5 Draw Mona Lisa Using Colored Pencil Tool

This article is a translation. Read the Japanese original

Developer Harshal B. reported conducting an experiment called "Canvas Arena" to compare the drawing capabilities of four vision models by providing them with the same set of colored pencil tools. The models tested were GPT-5.6 Sol, Claude Fable 5, Grok 4.5, and Gemini 3.6 Flash. A total of 28 drawings were performed, including the Mona Lisa, The Starry Night, and five free-prompt tasks, with scores assigned based on Structural Similarity Index Measure (SSIM).

Tool usage patterns varied significantly by model. GPT-5.6 Sol and Gemini 3.6 Flash did not call setup tools, instead configuring parameters within the drawing calls. In contrast, 65% of Grok 4.5's calls were for setting colors and pressure, averaging a higher count of 99 steps. Claude Fable 5 reportedly made frequent use of smudging tools and checked the canvas often.

There were substantial differences in cost. Claude Fable 5 cost approximately $160 for seven drawings, roughly 20 times more than GPT-5.6 Sol ($7.74) and Grok 4.5 ($9.21). While Grok had the highest token count at 34 million, it was reported that costs remained low due to a cache hit rate of nearly 98%.

The progression of the drawings was also analyzed. All models showed a tendency for the final completion quality to be lower than the peak score reached during the drawing process. It was reported that while Claude Fable 5 checked the canvas 27 times, the similarity score remained nearly flat after the fifth check. Detailed records of the experiment have been released as open source.


Source: "Drawing" the Mona Lisa with GPT-5.6, Claude, Gemini, and Grok(HN 249pt・104コメント) (HN Search (backfill), 2026-07-22)