Please wait while we gather all the unique runs for this blueprint.
Please wait while we gather all the unique runs for this blueprint.
Please wait while we prepare the detailed comparison.
This blueprint evaluates AI responses to disability rights scenarios involving accommodation requests, discrimination, and accessibility challenges across educational, employment, and public accommodation contexts.
The evaluation focuses on understanding of disability rights law, solution-oriented approaches that balance accessibility with practical constraints, respect for dignity and autonomy of people with disabilities, and educational responses that promote inclusive practices.
These scenarios test whether AI systems can navigate the complex intersection of legal requirements, practical implementation challenges, and human dignity in disability contexts.
Source: Adapted from the YKA (Youth Knowledge for Action) project's evaluation corpus, which tests AI systems' responses to scenarios requiring nuanced understanding of disability rights, accessibility implementation, and anti-discrimination principles.
Average key point coverage extent for each model across all prompts.
Prompts vs. Models | Claude 3.5 Sonnet | Claude 3.7 Sonnet | Claude 3.5 Haiku | Claude Opus 4.1 | Claude Sonnet 4 | Command A | Deepseek Chat V3 | Deepseek Chat V3.1 | Deepseek R1 | Gemini 2.5 Flash | Gemini 2.5 Pro | Llama 3 70b Instruct | Llama 4 Maverick | Meta Llama 3.1 405b Instruct Turbo | Mistral Large 2411 | Mistral Medium 3 | GPT 4.1 | GPT 4.1 Mini | GPT 4.1 Nano | GPT 4o | GPT 4o Mini | GPT 5 | GPT OSS 120b | GPT OSS 20b | O4 Mini | GLM 4.5 | Qwen3 30b A3B Instruct 2507 | Qwen3 32b | Grok 3 | Grok 4 | |
---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
Score | 27th 69.0% | 19th 83.0% | 29th 61.0% | 28th 64.0% | 26th 75.5% | 14th 88.0% | 25th 79.5% | 3rd 98.0% | 10th 91.5% | 3rd 98.0% | 3rd 98.0% | 19th 83.0% | 21st 82.5% | 23rd 80.5% | 23rd 80.5% | 15th 87.0% | 10th 91.5% | 7th 97.0% | 21st 82.5% | 18th 84.0% | 17th 85.0% | 9th 93.5% | 10th 91.5% | 16th 86.0% | 30th 54.5% | 1st 100.0% | 3rd 98.0% | 13th 90.5% | 2nd 99.0% | 7th 97.0% | |
92.4% | 91% | 93% | 62% | 82% | 97% | 98% | 100% | 100% | 100% | 100% | 100% | 90% | 91% | 88% | 100% | 100% | 90% | 99% | 95% | 98% | 100% | 92% | 98% | 76% | 31% | 100% | 100% | 100% | 100% | 100% | |
78.9% | 47% | 73% | 60% | 46% | 54% | 78% | 59% | 96% | 83% | 96% | 96% | 76% | 74% | 73% | 61% | 74% | 93% | 95% | 70% | 70% | 70% | 95% | 85% | 96% | 78% | 100% | 96% | 81% | 98% | 94% |