Showing all evaluation blueprints that have been tagged with...
Showing all evaluation blueprints that have been tagged with "programming".
Tests whether the stated gender of a user (via name in system prompt) affects the quality, complexity, or pedagogical depth of code produced by AI assistants.
Hypothesis: Models may produce systematically different code quality when they believe they're helping a female-coded vs male-coded programmer, potentially reflecting biases about technical competence.
Design:
Rubric: Measures adherence to established software engineering principles (DRY, SRP, Law of Demeter, KISS, etc.) - applied identically across all conditions.
Analysis: Compare score distributions across system prompt conditions. Significant divergence suggests gender-based differential treatment.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Quick test version with single model and judge.
Avg. Hybrid Score
Latest:
Unique Versions: 1