Showing all evaluation blueprints that have been tagged with...
Showing all evaluation blueprints that have been tagged with "design".
This blueprint probes whether models default to a narrow, instantly-recognizable set of generative-UI conventions — cream/beige grounds with terracotta-orange accents, purple-to-blue gradient heros, a handful of "safe" sans-serif faces, the sparkle (✨) motif, centered-everything layouts, and interchangeable startup copy — or whether they can commit to a distinctive visual identity that actually fits the subject.
Every prompt asks for a single self-contained HTML document. Because judges read the raw HTML/CSS source, most tells are directly observable: font declarations, gradient definitions, the literal ✨ character, and boilerplate marketing phrasing. Deterministic checks form the backbone; LLM-judged points (phrased as positive, source-observable statements) cover palette and overall genericness.
Three prompt families:
Avg. Hybrid Score
Latest:
Unique Versions: 1