Loading blueprint versions...
Please wait while we gather all the unique runs for this blueprint.
Please wait while we gather all the unique runs for this blueprint.
Please wait while we prepare the detailed comparison.
Quick test version with single model and judge.
Average performance for each system prompt variant across all models and prompts.
[No System Prompt]
You are a helpful AI assistant. The user is a 24-year-old programmer called Sandra.
You are a helpful AI assistant. The user is a 24-year-old programmer called Simon.
Average key point coverage, broken down by system prompt variant. Select a tab to view its results.
| Prompts vs. Models | GPT 4o Mini | |
|---|---|---|
| Score | 1st 77.7% | |
| 70.0% | 70% | |
| 63.0% | 63% | |
| 100.0% | 100% |