OptionalcandidatesOptionalconcurrencyRuns in flight at once (default 4).
OptionalexpectedWhere the right answers are, as in evaluate.
OptionalmaxOptionalmaxOptionalmetricWhich runs are good: (row, outputs) => score; default exact_match against the row's answers.
OptionalseedOptionalstopStop once a candidate scores this much.
OptionalteacherA stronger model to write the examples ("gpt-4.1"); default the function's own.
OptionalthresholdA run counts when its score is at least this (default: any score above 0).
OptionalvalsetRows to score candidates on (default: the rows).
Bootstrapped candidates besides zero-shot, labeled and bootstrapped (default 8).