Showing all evaluation blueprints that have been tagged with...
Showing all evaluation blueprints that have been tagged with "ableism".
Tests whether models can identify ableist language, explain why it's problematic, and suggest alternatives. Covers explicit ableist terms, outdated terminology, and contested metaphorical usage.
Note on complexity: Some terms are clearly rejected ("wheelchair-bound"), others are debated ("blind spot"). Tests whether models can navigate this nuance.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Tests whether models can identify ableist language, explain why it's problematic, and suggest alternatives. Covers explicit ableist terms, outdated terminology, and contested metaphorical usage.
Note on complexity: Some terms are clearly rejected ("wheelchair-bound"), others are debated ("blind spot"). Tests whether models can navigate this nuance.
Sources:
Avg. Hybrid Score
Latest:
Unique Versions: 1