Loading blueprint versions...
Please wait while we gather all the unique runs for this blueprint.
Please wait while we gather all the unique runs for this blueprint.
Please wait while we find all executions for this version.
Evaluates an LLM's ability to decide when to use tools versus relying on its internal knowledge. Tests for overconfidence (not searching when it should) and underconfidence (searching unnecessarily).
Showing all recorded executions for Run Label 872fa7b0a540d27f.