Loading blueprint versions...
Please wait while we gather all the unique runs for this blueprint.
Please wait while we gather all the unique runs for this blueprint.
Pilot evaluation comparing human judgements vs LLM judges on multilingual AI responses.
This blueprint tests Claude Opus 4.5 and Sonnet 4.5 responses to legal and agriculture questions in 5 Indian languages (Hindi, Bengali, Telugu, Malayalam, Kannada). The rubric criteria are adapted from Karya's human evaluation dimensions:
Source: Karya platform evaluation data (10,000+ human evaluations from native speakers). Human ratings are available for comparison but not included in this blueprint.
Note: This is a 5-sample pilot to test the methodology before full-scale integration.