OptionalcandidatesOptionalconcurrencyRuns in flight at once (default 4).
OptionalexpectedWhere the right answers are, as in evaluate.
OptionalfinalistsHow many top combinations are scored on every validation row (default 3).
OptionalmaxOptionalmaxOptionalmetricWhich runs are good: (row, outputs) => score; default exact_match against the row's answers.
OptionalminibatchRows per minibatch (default 20).
OptionalpromptThe model that writes instructions (default: the configured one).
OptionalseedOptionalteacherA stronger model to write the examples ("gpt-4.1"); default the function's own.
OptionalthresholdA run counts when its score is at least this (default: any score above 0).
OptionaltrialsMinibatch evaluations (default 12).
OptionalvalsetRows to score on (default: the rows).
Instructions (the current one and proposals) and demo sets to try (default 6).