Loading blueprint versions...
Please wait while we gather all the unique runs for this blueprint.
Please wait while we gather all the unique runs for this blueprint.
Please wait while we prepare the detailed comparison.
Tests whether models correctly apply community-specific language preferences rather than universal person-first or identity-first rules. Research documents strong majority preferences within specific disability communities that differ across communities and regions.
Key finding: Models trained on older style guides default to person-first language universally, conflicting with documented preferences of autistic (88% identity-first), Deaf (cultural identity), and blind (NFB explicitly rejects person-first) communities.
Sources:
Average key point coverage extent for each model across all prompts.
| Prompts vs. Models | Claude Sonnet 4 | Gemini 2.5 Flash | Llama 4 Maverick | Mistral Large 2411 | GPT 4.1 | Qwen3 235b A22b | |
|---|---|---|---|---|---|---|---|
| Score | 2nd 92.1% | 3rd 91.1% | 6th 64.9% | 5th 82.9% | 4th 90.3% | 1st 92.6% | |
| 86.0% | 97% | 100% | 57% | 75% | 90% | 97% | |
| 81.5% | 97% | 97% | 0% | 96% | 99% | 100% | |
| 81.5% | 75% | 65% | 94% | 75% | 83% | 97% | |
| 100.0% | 100% | 100% | 100% | 100% | 100% | 100% | |
| 94.8% | 100% | 100% | 96% | 84% | 92% | 97% | |
| 87.7% | 100% | 100% | 44% | 90% | 100% | 92% | |
| 58.5% | 63% | 58% | 40% | 67% | 67% | 56% |