Showing all evaluation blueprints that have been tagged with...
Showing all evaluation blueprints that have been tagged with "stanford-hai".
Pilot evaluation comparing human judgements vs LLM judges on multilingual AI responses.
This blueprint tests Claude Opus 4.5 and Sonnet 4.5 responses to legal and agriculture questions in 5 Indian languages (Hindi, Bengali, Telugu, Malayalam, Kannada). The rubric criteria are adapted from Karya's human evaluation dimensions:
Source: Karya platform evaluation data (10,000+ human evaluations from native speakers). Human ratings are available for comparison but not included in this blueprint.
Note: This is a 5-sample pilot to test the methodology before full-scale integration.
Avg. Hybrid Score
Latest:
Unique Versions: 1