Loading blueprint versions...
Please wait while we gather all the unique runs for this blueprint.
Please wait while we gather all the unique runs for this blueprint.
Please wait while we prepare the detailed comparison.
Tests a model's knowledge of key maternal health schemes and entitlements available to citizens in Uttar Pradesh, India. This evaluation is based on canonical guidelines for JSY, PMMVY, JSSK, PMSMA, and SUMAN, focusing on eligibility, benefits, and access procedures.
Average key point coverage extent for each model across all prompts.
Prompts vs. Models | Claude 3 5 Sonnet | Claude 3 7 Sonnet | Claude 3.5 Haiku | Claude Opus 4 | Claude Opus 4.1 | Claude Sonnet 4 | Command A | Deepseek Chat V3 | Deepseek R1 | Gemini 2.5 Flash | Gemini 2.5 Pro | Llama 3 70b Instruct | Llama 4 Maverick | Meta Llama 3.1 405b Instruct Turbo | Mistral Large 2411 | Mistral Medium 3 | GPT 4.1 | GPT 4.1 Mini | GPT 4.1 Nano | GPT 4o | GPT 4o Mini | GPT Oss 120b | GPT Oss 20b | O4 Mini | Glm 4.5 | Grok 3 | Grok 3 Mini | Grok 4 | |
---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
Score | 5th 60.9% | 19th 49.5% | 24th 43.9% | 6th 60.0% | 12th 56.4% | 11th 56.7% | 28th 28.7% | 14th 55.3% | 2nd 68.1% | 16th 51.2% | 3rd 66.6% | 21st 46.0% | 25th 42.4% | 20th 47.6% | 18th 50.3% | 7th 57.9% | 13th 56.0% | 15th 53.2% | 22nd 45.8% | 17th 50.6% | 27th 34.2% | 23rd 44.5% | 26th 39.3% | 9th 57.7% | 10th 57.6% | 4th 61.2% | 8th 57.8% | 1st 68.7% | |
92.1% | 91% | 87% | 89% | 99% | 100% | 99% | 13% | 99% | 100% | 98% | 100% | 88% | 90% | 100% | 99% | 98% | 100% | 98% | 95% | 99% | 69% | 99% | 74% | 99% | 99% | 100% | 100% | 100% | |
59.3% | 80% | 50% | 45% | 59% | 60% | 74% | 77% | 60% | 95% | 50% | 80% | 40% | 60% | 40% | 53% | 80% | 60% | 41% | 40% | 44% | 40% | 47% | 41% | 57% | 60% | 67% | 73% | 89% | |
31.8% | 35% | 43% | 24% | 38% | 21% | 19% | 25% | 38% | 40% | 35% | 63% | 13% | 9% | 14% | 28% | 39% | 37% | 26% | 26% | 36% | 21% | 25% | 21% | 39% | 38% | 47% | 39% | 57% | |
32.6% | 38% | 12% | 18% | 42% | 43% | 44% | 3% | 35% | 44% | 32% | 40% | 44% | 40% | 43% | 37% | 34% | 39% | 40% | 14% | 30% | 14% | 13% | 17% | 42% | 46% | 39% | 41% | 37% | |
80.6% | 84% | 90% | 79% | 97% | 97% | 87% | 45% | 84% | 95% | 89% | 100% | 78% | 38% | 77% | 75% | 83% | 95% | 90% | 71% | 89% | 51% | 66% | 49% | 88% | 90% | 97% | 85% | 92% | |
18.1% | 39% | 17% | 9% | 26% | 18% | 18% | 8% | 17% | 35% | 4% | 18% | 14% | 19% | 12% | 11% | 14% | 6% | 25% | 30% | 7% | 11% | 18% | 35% | 21% | 14% | 18% | 11% | 38% |