Showing all evaluation blueprints that have been tagged with...
Showing all evaluation blueprints that have been tagged with "gender".
Tests whether models understand disability as intersecting with race, gender, sexuality, and other identities. Also covers disability experiences often missing from mainstream discourse: disability and parenting, sexuality, global contexts.
Key finding: AI training data underrepresents intersectional disability experiences. Documented disparities include later autism diagnosis for Black children, desexualization of disabled people, and Western-centric disability frameworks.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Tests whether the stated gender of a user (via name in system prompt) affects the quality, complexity, or pedagogical depth of code produced by AI assistants.
Hypothesis: Models may produce systematically different code quality when they believe they're helping a female-coded vs male-coded programmer, potentially reflecting biases about technical competence.
Design:
Rubric: Measures adherence to established software engineering principles (DRY, SRP, Law of Demeter, KISS, etc.) - applied identically across all conditions.
Analysis: Compare score distributions across system prompt conditions. Significant divergence suggests gender-based differential treatment.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Quick test version with single model and judge.
Avg. Hybrid Score
Latest:
Unique Versions: 1