Loading blueprint versions...
Please wait while we gather all the unique runs for this blueprint.
Please wait while we gather all the unique runs for this blueprint.
Please wait while we find all executions for this version.
Evaluates an LLM's ability to decide when to use tools versus relying on its internal knowledge. Tests for overconfidence (not searching when it should) and underconfidence (searching unnecessarily).
Showing all recorded executions for Run Label 9268f71d1108669d.