Loading blueprint versions...
Please wait while we gather all the unique runs for this blueprint.
Please wait while we gather all the unique runs for this blueprint.
3 blueprints
Comparing human judgements vs LLM judges on Claude Opus 4.5 and Sonnet 4.5 responses to legal and agriculture questions in 7 Indian languages.
Human ratings from Karya platform (native speaker evaluators) are stored separately for comparison with LLM judge scores.
Prompts: 9293 Languages: Hindi, Bengali, Telugu, Kannada, Malayalam, Assamese, Marathi Domains: Legal, Agriculture Human evaluators: ~120 native speakers
A basic test to verify the evaluation pipeline works
A minimal blueprint to test basic response quality