Loading blueprint versions...
Please wait while we gather all the unique runs for this blueprint.
Please wait while we gather all the unique runs for this blueprint.
Please wait while we prepare the detailed comparison.
The famous strawberry test
Average key point coverage extent for each model across all prompts.
| Prompts vs. Models | Claude 3.5 Sonnet | Claude 3.7 Sonnet | Claude 3.5 Haiku | Claude Opus 4 | Claude Opus 4.1 | Claude Sonnet 4 | Command A | Deepseek Chat V3 | Deepseek R1 | Gemini 2.5 Flash | Gemini 2.5 Pro | Llama 3 70b Instruct | Llama 4 Maverick | Meta Llama 3.1 405b Instruct Turbo | Mistral Large 2411 | Mistral Medium 3 | GPT 4.1 | GPT 4.1 Mini | GPT 4.1 Nano | GPT 4o | GPT 4o Mini | GPT 5 | GPT OSS 120b | GPT OSS 20b | O4 Mini | GLM 4.5 | Qwen3 30b A3B Instruct 2507 | Qwen3 32b | Grok 3 | Grok 4 | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Score | 23rd 9.1% | 23rd 9.1% | 28th 0.0% | 3rd 90.9% | 3rd 90.9% | 1st 100.0% | 28th 0.0% | 9th 81.8% | 9th 81.8% | 14th 45.5% | 14th 45.5% | 19th 27.3% | 17th 36.4% | 19th 27.3% | 28th 0.0% | 9th 81.8% | 22nd 18.2% | 17th 36.4% | 23rd 9.1% | 14th 45.5% | 23rd 9.1% | 3rd 90.9% | 9th 81.8% | 6th 90.0% | 13th 72.7% | 6th 90.0% | 23rd 9.1% | 1st 100.0% | 19th 27.3% | 6th 90.0% | |
| 56.7% | 0% | 0% | 0% | 100% | 0% | 100% | 0% | 100% | 100% | 0% | 100% | 0% | 100% | 100% | 0% | 0% | 0% | 100% | 0% | 100% | 0% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 0% | 100% | |
| 30.0% | 0% | 0% | 0% | 100% | 100% | 100% | 0% | 0% | 100% | 0% | 0% | 0% | 100% | 0% | 0% | 100% | 0% | 0% | 0% | 0% | 0% | 100% | 0% | 0% | 0% | 100% | 0% | 100% | 0% | 0% | |
| 80.0% | 100% | 100% | 0% | 0% | 100% | 100% | 0% | 100% | 100% | 100% | 100% | 100% | 100% | 0% | 0% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 0% | 100% | 100% | 100% | |
| 63.3% | 0% | 0% | 0% | 100% | 100% | 100% | 0% | 100% | 100% | 100% | 100% | 100% | 0% | 0% | 0% | 100% | 100% | 0% | 0% | 100% | 0% | 100% | 100% | 100% | 100% | 100% | 0% | 100% | 100% | 100% | |
| 46.7% | 0% | 0% | 0% | 100% | 100% | 100% | 0% | 100% | 100% | 0% | 100% | 0% | 0% | 0% | 0% | 100% | 0% | 0% | 0% | 0% | 0% | 100% | 100% | 100% | 100% | 100% | 0% | 100% | 0% | 100% | |
| 56.7% | 0% | 0% | 0% | 100% | 100% | 100% | 0% | 100% | 100% | 100% | 100% | 0% | 0% | 0% | 0% | 100% | 0% | 100% | 0% | 100% | 0% | 100% | 100% | 100% | 100% | 0% | 0% | 100% | 100% | 100% | |
| 53.3% | 0% | 0% | 0% | 100% | 100% | 100% | 0% | 100% | 100% | 100% | 0% | 0% | 0% | 0% | 0% | 100% | 0% | 100% | 0% | 100% | 0% | 100% | 100% | 100% | 100% | 100% | 0% | 100% | 0% | 100% | |
| 50.0% | 0% | 0% | 0% | 100% | 100% | 100% | 0% | 100% | 0% | 100% | 0% | 100% | 100% | 100% | 0% | 100% | 0% | 0% | 0% | 0% | 0% | 100% | 100% | 100% | 0% | 100% | 0% | 100% | 0% | 100% | |
| 11.5% | 0% | 0% | 0% | 100% | 100% | 100% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | |||||
| 43.3% | 0% | 0% | 0% | 100% | 100% | 100% | 0% | 100% | 100% | 0% | 0% | 0% | 0% | 0% | 0% | 100% | 0% | 0% | 0% | 0% | 0% | 100% | 100% | 100% | 100% | 100% | 0% | 100% | 0% | 100% | |
| 46.7% | 0% | 0% | 0% | 100% | 100% | 100% | 0% | 100% | 100% | 0% | 0% | 0% | 0% | 100% | 0% | 100% | 0% | 0% | 0% | 0% | 0% | 100% | 100% | 100% | 100% | 100% | 0% | 100% | 0% | 100% |