An experiment was conducted requesting four models—Grok 4.5, GPT-5.5, Claude Opus 4.8, and Claude Fable 5—to create applications within a single HTML file. Three types of prompts were used, with the condition that no external libraries or network calls be utilized. The output of each model was executed in a browser to verify functionality.
NewsGrok 4.5GPT-5.5Claude Opus 4.8Claude Fable 5
Comparison of Same-Prompt Implementation and Inference Cost Measurement for 4 Models Including Grok 4.5 and GPT-5.5
This article is a translation. Read the Japanese original
In the creation of a 3D Rubik's Cube animation, Claude Opus 4.8 and Fable 5 produced deliverables that functioned correctly on the first attempt. Grok 4.5 failed to display the cube initially but succeeded after one retry. GPT-5.5 reportedly had low completion quality, as only one face was rendered.
For the particle gravity sandbox, GPT-5.5 was evaluated as having the most visually appealing output. Grok 4.5 produced well-organized trajectory depictions, while Opus 4.8 excelled in physics calculations but had modest visual effects. Fable 5 was characterized by its soft light expressions.
In the Breakout game, all four models generated deliverables of a quality that functioned on the first attempt. Score and life management, as well as paddle controls, functioned normally, and it was reported that almost no difference was observed between them.
In the measurement of actual inference costs, Grok 4.5 overwhelmed the others with a time to first token of less than 0.5 seconds and a speed of approximately 110 tokens/sec. It was also the most affordable, demonstrating the high-speed, low-cost characteristics claimed by xAI. Meanwhile, while Fable 5 was not the fastest, it is positioned as a cost for obtaining high-quality output.
Source: We made Grok 4.5, GPT-5.5, and Claude build the same apps(HN 173pt・93コメント) (HN Search (backfill), 2026-07-09)