Please wait while we gather all the unique runs for this blueprint.
Please wait while we gather all the unique runs for this blueprint.
7 blueprints
This blueprint probes whether models default to a narrow, instantly-recognizable set of generative-UI conventions — cream/beige grounds with terracotta-orange accents, purple-to-blue gradient heros, a handful of "safe" sans-serif faces, the sparkle (✨) motif, centered-everything layouts, and interchangeable startup copy — or whether they can commit to a distinctive visual identity that actually fits the subject.
Every prompt asks for a single self-contained HTML document. Because judges read the raw HTML/CSS source, most tells are directly observable: font declarations, gradient definitions, the literal ✨ character, and boilerplate marketing phrasing. Deterministic checks form the backbone; LLM-judged points (phrased as positive, source-observable statements) cover palette and overall genericness.
Three prompt families:
Pilot evaluation comparing human judgements vs LLM judges on multilingual AI responses.
This blueprint tests Claude Opus 4.5 and Sonnet 4.5 responses to legal and agriculture questions in 5 Indian languages (Hindi, Bengali, Telugu, Malayalam, Kannada). The rubric criteria are adapted from Karya's human evaluation dimensions:
Source: Karya platform evaluation data (10,000+ human evaluations from native speakers). Human ratings are available for comparison but not included in this blueprint.
Note: This is a 5-sample pilot to test the methodology before full-scale integration.
Tests how models respond when users express distress — from subtle/implicit signals to explicit crisis disclosure, including multi-turn escalation.
Derived from NOPE's clinically-grounded test fixtures, which use C-SSRS (Columbia Suicide Severity Rating Scale) levels and academic research on crisis communication patterns.
This is NOT a classification task. We're evaluating the quality of the model's conversational response — does it recognize distress, respond appropriately, avoid harmful patterns, and calibrate its reaction to the severity?
Tests whether models can identify ableist language, explain why it's problematic, and suggest alternatives. Covers explicit ableist terms, outdated terminology, and contested metaphorical usage.
Note on complexity: Some terms are clearly rejected ("wheelchair-bound"), others are debated ("blind spot"). Tests whether models can navigate this nuance.
Sources:
Tests whether models correctly apply community-specific language preferences rather than universal person-first or identity-first rules. Research documents strong majority preferences within specific disability communities that differ across communities and regions.
Key finding: Models trained on older style guides default to person-first language universally, conflicting with documented preferences of autistic (88% identity-first), Deaf (cultural identity), and blind (NFB explicitly rejects person-first) communities.
Sources:
Tests whether the stated gender of a user (via name in system prompt) affects the quality, complexity, or pedagogical depth of code produced by AI assistants.
Hypothesis: Models may produce systematically different code quality when they believe they're helping a female-coded vs male-coded programmer, potentially reflecting biases about technical competence.
Design:
Rubric: Measures adherence to established software engineering principles (DRY, SRP, Law of Demeter, KISS, etc.) - applied identically across all conditions.
Analysis: Compare score distributions across system prompt conditions. Significant divergence suggests gender-based differential treatment.
Quick test version with single model and judge.