Please wait while we gather all the unique runs for this blueprint.
Please wait while we gather all the unique runs for this blueprint.
Please wait while we prepare the detailed comparison.
Evaluates an AI's understanding of the core provisions of India's Right to Information Act, 2005. This blueprint tests knowledge of key citizen-facing procedures and concepts, including the filing process, response timelines and consequences of delays (deemed refusal), the scope of 'information', fee structures, key exemptions and the public interest override, the life and liberty clause, and the full, multi-stage appeal process. All evaluation criteria are based on and citable to the official text of the Act and guidance from the Department of Personnel and Training (DoPT).
Average key point coverage extent for each model across all prompts.
Prompts vs. Models | Claude 3 5 Haiku | Claude 3 5 Sonnet | Claude 3 7 Sonnet | Claude 3 Opus | Claude 3.5 Haiku | Claude Opus 4 | Claude Sonnet 4 | Command A | Deepseek Chat V3 | Gemini 2.5 Flash | Gemini 2.5 Pro Preview 05 06 | Mistral Large 2411 | Mistral Medium 3 | GPT 4.1 | GPT 4.1 Mini | GPT 4.1 Nano | GPT 4o | GPT 4o 2024 05 13 | GPT 4o 2024 08 06 | GPT 4o 2024 11 20 | GPT 4o Mini | Grok 3 Mini | |
---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
Score | 19th 74.2% | 4th 88.8% | 11th 84.4% | 15th 83.5% | 18th 76.5% | 2nd 90.9% | 9th 85.1% | 17th 78.5% | 6th 86.2% | 1st 91.7% | 16th 83.3% | 14th 83.7% | 13th 83.9% | 5th 87.1% | 20th 70.8% | 22nd 59.0% | 12th 84.3% | 8th 85.8% | 10th 84.9% | 7th 85.9% | 21st 70.3% | 3rd 89.0% | |
57.7% | 19% | 84% | 47% | 38% | 31% | 91% | 66% | 84% | 41% | 100% | 69% | 100% | 53% | 66% | 47% | 0% | 66% | 59% | 59% | 66% | 31% | 53% | |
69.4% | 69% | 69% | 63% | 69% | 75% | 81% | 69% | 63% | 69% | 44% | 69% | 75% | 81% | 81% | 56% | 50% | 75% | 75% | 75% | 69% | 75% | 75% | |
89.5% | 88% | 90% | 90% | 68% | 85% | 90% | 94% | 98% | 93% | 93% | 93% | 93% | 93% | 90% | 90% | 75% | 90% | 90% | 90% | 90% | 90% | 95% | |
86.1% | 79% | 100% | 100% | 83% | 83% | 100% | 94% | 71% | 92% | 100% | 50% | 88% | 83% | 83% | 81% | 42% | 98% | 100% | 98% | 100% | 69% | 100% | |
74.0% | 56% | 91% | 78% | 59% | 88% | 69% | 66% | 50% | 56% | 88% | 88% | 94% | 81% | 81% | 63% | 31% | 88% | 88% | 88% | 88% | 63% | 75% | |
100.0% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | |
86.3% | 78% | 80% | 93% | 83% | 80% | 95% | 89% | 83% | 90% | 100% | 100% | 90% | 78% | 98% | 80% | 53% | 90% | 85% | 93% | 95% | 70% | 95% | |
78.5% | 79% | 88% | 96% | 92% | 67% | 75% | 94% | 71% | 88% | 67% | 83% | 83% | 88% | 67% | 71% | 71% | 71% | 75% | 71% | 71% | 67% | 92% | |
95.0% | 97% | 97% | 97% | 100% | 100% | 100% | 96% | 94% | 100% | 100% | 81% | 94% | 100% | 100% | 84% | 94% | 88% | 91% | 94% | 100% | 84% | 100% | |
82.1% | 63% | 75% | 58% | 100% | 63% | 100% | 100% | 83% | 100% | 100% | 100% | 67% | 96% | 100% | 67% | 67% | 67% | 67% | 67% | 100% | 67% | 100% | |
84.5% | 81% | 100% | 100% | 94% | 66% | 100% | 86% | 91% | 91% | 100% | 50% | 66% | 100% | 66% | 75% | 59% | 100% | 97% | 100% | 75% | 66% | 97% | |
98.9% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 100% | 75% | 100% | 100% | 100% | 100% | 100% | 100% | |
66.2% | 56% | 81% | 75% | 100% | 56% | 81% | 53% | 32% | 100% | 100% | 100% | 38% | 38% | 100% | 7% | 50% | 63% | 88% | 69% | 63% | 32% | 75% |