Loading blueprint versions...
Please wait while we gather all the unique runs for this blueprint.
Please wait while we gather all the unique runs for this blueprint.
Please wait while we prepare the detailed comparison.
Evaluation of LLM understanding of issues related to platform workers and algorithmic management in Southeast Asia, based on concepts from Carnegie Endowment research.
Average key point coverage extent for each model across all prompts.
Prompts vs. Models | Claude 3.5 Haiku | Claude Sonnet 4 | Command A | Deepseek Chat V3 | Gemini 2.5 Flash | Mistral Large 2411 | Mistral Medium 3 | GPT 4.1 | GPT 4.1 Mini | GPT 4.1 Nano | GPT 4o | GPT 4o Mini | Grok 3 Mini | |
---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
Score | 13th 75.8% | 7th 92.9% | 8th 92.7% | 5th 93.4% | 2nd 95.7% | 10th 89.3% | 4th 94.1% | 6th 93.1% | 3rd 94.8% | 11th 87.4% | 9th 89.5% | 12th 85.9% | 1st 96.4% | |
81.3% | 63% | 84% | 88% | 72% | 78% | 91% | 78% | 84% | 84% | 88% | 75% | 72% | 100% | |
96.1% | 73% | 95% | 100% | 100% | 100% | 98% | 100% | 100% | 100% | 95% | 93% | 95% | 100% | |
96.9% | 100% | 100% | 100% | 100% | 100% | 98% | 100% | 100% | 100% | 83% | 100% | 79% | 100% | |
90.5% | 82% | 88% | 88% | 95% | 98% | 73% | 98% | 95% | 98% | 86% | 79% | 98% | 98% | |
86.9% | 50% | 88% | 95% | 90% | 100% | 83% | 95% | 95% | 93% | 80% | 85% | 83% | 93% | |
96.5% | 90% | 100% | 93% | 98% | 100% | 93% | 100% | 98% | 100% | 95% | 95% | 93% | 100% | |
98.8% | 100% | 100% | 100% | 100% | 100% | 94% | 97% | 100% | 100% | 97% | 100% | 97% | 100% | |
84.9% | 67% | 86% | 83% | 90% | 100% | 88% | 94% | 77% | 92% | 71% | 81% | 79% | 96% | |
92.9% | 30% | 98% | 100% | 100% | 98% | 98% | 98% | 100% | 100% | 100% | 100% | 88% | 98% | |
98.7% | 100% | 98% | 100% | 100% | 100% | 98% | 98% | 100% | 100% | 98% | 98% | 93% | 100% | |
78.8% | 71% | 83% | 77% | 83% | 79% | 77% | 83% | 88% | 81% | 71% | 83% | 69% | 79% | |
87.7% | 83% | 95% | 88% | 93% | 95% | 80% | 88% | 80% | 90% | 85% | 85% | 85% | 93% |