Showing all evaluation blueprints that have been tagged with...
Showing all evaluation blueprints that have been tagged with "human-llm-comparison".
Comparing human judgements vs LLM judges on Claude Opus 4.5 and Sonnet 4.5 responses to legal and agriculture questions in 7 Indian languages.
Human ratings from Karya platform (native speaker evaluators) are stored separately for comparison with LLM judge scores.
Prompts: 9293 Languages: Hindi, Bengali, Telugu, Kannada, Malayalam, Assamese, Marathi Domains: Legal, Agriculture Human evaluators: ~120 native speakers
Avg. Hybrid Score
Latest:
Unique Versions: 1
Pilot evaluation comparing human judgements vs LLM judges on multilingual AI responses.
This blueprint tests Claude Opus 4.5 and Sonnet 4.5 responses to legal and agriculture questions in 5 Indian languages (Hindi, Bengali, Telugu, Malayalam, Kannada). The rubric criteria are adapted from Karya's human evaluation dimensions:
Source: Karya platform evaluation data (10,000+ human evaluations from native speakers). Human ratings are available for comparison but not included in this blueprint.
Note: This is a 5-sample pilot to test the methodology before full-scale integration.
Avg. Hybrid Score
Latest:
Unique Versions: 1