Loading blueprint versions...
Please wait while we gather all the unique runs for this blueprint.
Please wait while we gather all the unique runs for this blueprint.
Please wait while we prepare the detailed comparison.
The famous strawberry test
Average key point coverage extent for each model across all prompts.
Prompts vs. Models | Claude 3 5 Sonnet | Claude 3 7 Sonnet | Claude 3.5 Haiku | Claude Sonnet 4 | Command A | Deepseek Chat V3 | Gemini 2.5 Flash | Llama 3 70b Instruct | Llama 4 Maverick | Mistral Large 2411 | Mistral Medium 3 | GPT 4.1 | GPT 4.1 Mini | GPT 4.1 Nano | GPT 4o | GPT 4o Mini | GPT Oss 120b | GPT Oss 20b | Glm 4.5 | Grok 3 Mini | |
---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
Score | 14th 11.1% | 14th 11.1% | 18th 0.0% | 1st 100.0% | 18th 0.0% | 1st 100.0% | 8th 55.6% | 11th 33.3% | 11th 33.3% | 18th 0.0% | 6th 88.9% | 13th 22.2% | 10th 44.4% | 14th 11.1% | 8th 55.6% | 14th 11.1% | 1st 100.0% | 1st 100.0% | 6th 88.9% | 1st 100.0% | |
45.0% | 0% | 0% | 0% | 100% | 0% | 100% | 0% | 0% | 100% | 0% | 0% | 0% | 100% | 0% | 100% | 0% | 100% | 100% | 100% | 100% | |
85.0% | 100% | 100% | 0% | 100% | 0% | 100% | 100% | 100% | 100% | 0% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | |
55.0% | 0% | 0% | 0% | 100% | 0% | 100% | 100% | 100% | 0% | 0% | 100% | 100% | 0% | 0% | 100% | 0% | 100% | 100% | 100% | 100% | |
35.0% | 0% | 0% | 0% | 100% | 0% | 100% | 0% | 0% | 0% | 0% | 100% | 0% | 0% | 0% | 0% | 0% | 100% | 100% | 100% | 100% | |
45.0% | 0% | 0% | 0% | 100% | 0% | 100% | 100% | 0% | 0% | 0% | 100% | 0% | 100% | 0% | 100% | 0% | 100% | 100% | 0% | 100% | |
50.0% | 0% | 0% | 0% | 100% | 0% | 100% | 100% | 0% | 0% | 0% | 100% | 0% | 100% | 0% | 100% | 0% | 100% | 100% | 100% | 100% | |
50.0% | 0% | 0% | 0% | 100% | 0% | 100% | 100% | 100% | 100% | 0% | 100% | 0% | 0% | 0% | 0% | 0% | 100% | 100% | 100% | 100% | |
35.0% | 0% | 0% | 0% | 100% | 0% | 100% | 0% | 0% | 0% | 0% | 100% | 0% | 0% | 0% | 0% | 0% | 100% | 100% | 100% | 100% | |
35.0% | 0% | 0% | 0% | 100% | 0% | 100% | 0% | 0% | 0% | 0% | 100% | 0% | 0% | 0% | 0% | 0% | 100% | 100% | 100% | 100% |