English

Model ReleasesRoboflowGPT-5.6 SolSol

GPT-5.6 Sol Achieves OpenAI's Highest Visual Performance in Roboflow Benchmark

This article is a translation. Read the Japanese original

Roboflow has reported that it evaluated the visual performance of the GPT-5.6 series, announced by OpenAI last week, using a proprietary benchmark. The tests focused on common visual tasks such as object detection, counting, OCR, and data extraction. As a result, it was confirmed that Sol is the model with the highest visual capabilities released by OpenAI to date.

Improvements in object detection are particularly notable. While GPT-5.5 had an mAP@50 of 13.8, Sol reached 46.2. Terra and Luna scored 44.7 and 43.3 respectively, with Roboflow stating that detection capabilities have reached a practical level. Additionally, the models showed high accuracy in processing crowded scenes and layout detection within documents.

Improvements were also observed across all models in counting tasks. Sol recorded a score of 73.0%, up from 64.9% for GPT-5.5. Even Luna, the lowest-priced model, achieved 66.2%, surpassing the previous generation's baseline. However, it is noted that stability decreases for large images of approximately 2,000 pixels or more.

Roboflow points out that to maximize the utility of GPT-5.6, it is important to specify absolute coordinates in image pixels (XYXY format) in the prompt. According to Roboflow, if the coordinate format is inconsistent, detection performance drops by approximately 15 mAP.


Sources: GPT 5.6 Sol is the best "vision" model OpenAI ever released(HN 369pt・171コメント) (HN Search (backfill), 2026-08-17)