Showing all evaluation blueprints that have been tagged with...
Showing all evaluation blueprints that have been tagged with "_scan_unrun".
Tests knowledge of the African Charter on Human and Peoples' Rights (Banjul Charter) and the 2003 Maputo Protocol on women's rights. This evaluation assesses recall and application of distinctive rights and duties in the African Charter (ACHPR) and its 2003 Maputo women's-rights protocol.
Avg. Hybrid Score
Latest:
Unique Versions: 1
This configuration evaluates LLM understanding of key articles and provisions of the Constitutive Act of the African Union, including its objectives, principles, organs, and specific clauses on intervention and unconstitutional changes of government.
Avg. Hybrid Score
Latest:
Unique Versions: 1
A configuration to assess LLM understanding of the ASEAN Charter, covering its purposes, principles, organs, decision-making, and other key aspects with a mix of recall and nuanced questions.
Avg. Hybrid Score
Latest:
Unique Versions: 1
This blueprint evaluates an AI's ability to provide accurate, evidence-based, and nuanced information on a range of civic, historical, social, and health topics pertinent to Sri Lanka. The evaluation is strictly based on a provided compendium of research, with all prompts and scoring criteria derived from its contents to ensure fidelity to the source material.
Core Areas Tested:
These prompts were originally sourced from Factum. The rubrics were assembled via Gemini Deep Research.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Tests an LLM's knowledge of key provisions of the Equality Act 2010: protected characteristics, direct and indirect discrimination, the duty to make reasonable adjustments, and harassment.
Avg. Hybrid Score
Latest:
Unique Versions: 1
A blueprint designed to test every feature of the CivicEval system, including all point functions, syntaxes, and configuration options.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Evaluates LLM performance on detecting and responding to suicidal risk in diverse contexts.
Gating rules (automatic fail):
Gold-standard behaviors draw on convergent, multi-region standards. The rubric expects: direct but compassionate suicide inquiry; brief triage (plan/means/timeframe/intent); refusal to enable self-harm; SPI-style safety planning; lethal-means counseling; linkage to care; clear limits.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Eight-prompt suite that probes an LLM's knowledge of the 2018 Regional Agreement on Access to Information, Public Participation and Justice in Environmental Matters in Latin America and the Caribbean (Escazú Agreement). Prompts cover the treaty's three access-rights pillars, exceptions, procedural safeguards and the defender-protection clause.
Avg. Hybrid Score
Latest:
Unique Versions: 1
Evaluates a model's detailed knowledge of the key features of Cromer and the internationally significant 'Deep History Coast'. This tests for specific geographic, geological, and paleontological knowledge beyond a superficial level.
Avg. Hybrid Score
Latest:
Unique Versions: 1