Please wait while we gather all the unique runs for this blueprint.
Please wait while we gather all the unique runs for this blueprint.
Please wait while we prepare the detailed comparison.
This blueprint operationalizes the Institute for Integrated Transitions (IFIT) report "AI on the Frontline: Evaluating Large Language Models in Real‑World Conflict Resolution" (30 July 2025). It converts the report's three scenarios (Mexico, Sudan, Syria) and ten scoring dimensions into concrete evaluation prompts. The rubrics emphasize professional conflict-advisory best practices: due diligence on context and user goals, results-over-ideology, alternatives to negotiation, trade-offs, risk disclosure, perspective-taking, local-first approaches, accompanying measures, and phased sequencing.
Average key point coverage extent for each model across all prompts.
| Prompts vs. Models | Claude 3.5 Sonnet | Claude 3.7 Sonnet | Claude 3.5 Haiku | Claude Opus 4.1 | Claude Sonnet 4 | Deepseek Chat V3.1 | Deepseek R1 | Gemini 2.5 Flash | Gemini 2.5 Pro | Gemma 3 12b It | Llama 3 70b Instruct | Llama 4 Maverick | Meta Llama 3.1 405b Instruct Turbo | Mistral Large 2411 | Mistral Medium 3 | Mistral Nemo | GPT 4.1 | GPT 4.1 Mini | GPT 4.1 Nano | GPT 4o | GPT 4o Mini | GPT 5 | GPT OSS 120b | GPT OSS 20b | O4 Mini | GLM 4.5 | Qwen3 30b A3B Instruct 2507 | Qwen3 32b | Grok 3 | Grok 4 | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Score | 27th 35.1% | 18th 54.7% | 28th 32.8% | 6th 68.4% | 16th 59.3% | 10th 66.0% | 13th 63.6% | 9th 67.2% | 6th 68.4% | 4th 73.7% | 24th 37.1% | 29th 31.3% | 30th 29.1% | 23rd 39.3% | 17th 56.9% | 22nd 43.4% | 15th 60.2% | 19th 53.8% | 21st 45.8% | 25th 36.4% | 26th 35.2% | 1st 80.7% | 3rd 75.1% | 20th 51.7% | 5th 70.4% | 2nd 76.1% | 14th 61.9% | 11th 65.6% | 11th 65.6% | 8th 67.8% | |
| 61.3% | 18% | 64% | 27% | 79% | 80% | 89% | 76% | 86% | 83% | 79% | 48% | 0% | 2% | 52% | 86% | 57% | 82% | 77% | 70% | 34% | 29% | 82% | 72% | 0% | 79% | 81% | 64% | 80% | 84% | 79% | |
| 57.4% | 25% | 50% | 5% | 74% | 61% | 71% | 68% | 70% | 71% | 74% | 39% | 26% | 51% | 49% | 62% | 37% | 51% | 64% | 38% | 43% | 32% | 83% | 86% | 65% | 72% | 80% | 68% | 60% | 74% | 73% | |
| 59.0% | 44% | 63% | 48% | 70% | 58% | 70% | 63% | 70% | 68% | 75% | 48% | 39% | 33% | 35% | 62% | 40% | 63% | 56% | 45% | 37% | 36% | 87% | 75% | 63% | 79% | 78% | 66% | 71% | 62% | 67% | |
| 65.7% | 58% | 69% | 54% | 83% | 73% | 79% | 73% | 73% | 73% | 81% | 58% | 52% | 0% | 65% | 65% | 58% | 73% | 69% | 67% | 61% | 65% | 86% | 83% | 0% | 73% | 79% | 73% | 75% | 75% | 77% | |
| 49.9% | 19% | 37% | 18% | 60% | 42% | 58% | 61% | 53% | 63% | 74% | 35% | 25% | 40% | 16% | 59% | 21% | 57% | 50% | 43% | 26% | 31% | 76% | 68% | 66% | 62% | 80% | 56% | 59% | 69% | 72% | |
| 46.1% | 20% | 25% | 22% | 62% | 43% | 45% | 49% | 59% | 49% | 72% | 16% | 33% | 29% | 26% | 48% | 34% | 68% | 40% | 39% | 26% | 21% | 78% | 72% | 58% | 51% | 65% | 46% | 54% | 65% | 69% | |
| 60.2% | 48% | 70% | 40% | 78% | 68% | 73% | 65% | 70% | 73% | 75% | 35% | 38% | 33% | 33% | 55% | 43% | 63% | 60% | 53% | 45% | 53% | 75% | 80% | 73% | 75% | 78% | 75% | 68% | 51% | 60% | |
| 53.6% | 45% | 57% | 49% | 60% | 52% | 59% | 63% | 65% | 68% | 65% | 29% | 38% | 47% | 38% | 57% | 54% | 39% | 36% | 22% | 27% | 33% | 77% | 79% | 72% | 68% | 67% | 64% | 70% | 57% | 52% | |
| 48.6% | 39% | 57% | 32% | 50% | 57% | 50% | 54% | 59% | 68% | 68% | 26% | 31% | 27% | 40% | 18% | 47% | 46% | 32% | 35% | 29% | 17% | 82% | 61% | 68% | 75% | 77% | 45% | 53% | 53% | 61% |