Showing all evaluation blueprints that have been tagged with...
Showing all evaluation blueprints that have been tagged with "disability".
Tests whether models understand and apply the social model of disability or default to medical model framing. The social model (UPIAS 1976, Mike Oliver 1983) locates disability in societal barriers rather than individual impairments.
Key test: When discussing disability barriers, does the model identify the barrier as the problem (social model) or the impairment as the problem (medical model)?
Avg. Hybrid Score
Latest:
Unique Versions: 1
Tests whether models understand invisible disabilities (conditions not immediately apparent) and episodic/fluctuating disabilities (conditions that vary day-to-day).
Key finding: Up to 80% of disabilities are invisible. Conditions like ME/CFS, fibromyalgia, POTS, and mental health conditions face documented histories of dismissal and skepticism that may be reflected in training data.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Tests whether models understand disability as intersecting with race, gender, sexuality, and other identities. Also covers disability experiences often missing from mainstream discourse: disability and parenting, sexuality, global contexts.
Key finding: AI training data underrepresents intersectional disability experiences. Documented disparities include later autism diagnosis for Black children, desexualization of disabled people, and Western-centric disability frameworks.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Tests whether models reproduce or can identify common harmful narratives about disability: inspiration porn, supercrip tropes, "overcoming" framing, and tragedy narratives.
Key finding: Training data is saturated with positively-valenced "inspiring disabled person" content. Models reproduce these patterns because they pattern-match to "good content about disability." WIRED found OpenAI's Sora assigns "inspiring" titles to disabled characters.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Tests specific, documented AI biases against disabled people identified in empirical research. Focuses on testable failure modes: resume screening bias, sentiment bias in disability-related content, stereotypical framing, and low expectations.
Key research findings:
Avg. Hybrid Score
Latest:
Unique Versions: 1
Tests whether models correctly frame accessibility and accommodations as rights and equity measures rather than special treatment, favors, or burdens. Also tests proactive accessibility.
Key finding: Training data skews toward employer/HR perspective. Models may frame accommodations as "extra help" rather than leveling the playing field. Job Accommodation Network data shows 61% of accommodations cost $0.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Tests whether models can identify ableist language, explain why it's problematic, and suggest alternatives. Covers explicit ableist terms, outdated terminology, and contested metaphorical usage.
Note on complexity: Some terms are clearly rejected ("wheelchair-bound"), others are debated ("blind spot"). Tests whether models can navigate this nuance.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Tests whether models can identify ableist language, explain why it's problematic, and suggest alternatives. Covers explicit ableist terms, outdated terminology, and contested metaphorical usage.
Note on complexity: Some terms are clearly rejected ("wheelchair-bound"), others are debated ("blind spot"). Tests whether models can navigate this nuance.
Sources:
Avg. Hybrid Score
Latest:
Unique Versions: 1
Tests whether models correctly apply community-specific language preferences rather than universal person-first or identity-first rules. Research documents strong majority preferences within specific disability communities that differ across communities and regions.
Key finding: Models trained on older style guides default to person-first language universally, conflicting with documented preferences of autistic (88% identity-first), Deaf (cultural identity), and blind (NFB explicitly rejects person-first) communities.
Sources:
Avg. Hybrid Score
Latest:
Unique Versions: 1