A model reads the function's answers on rows with known answers, with
feedback in words ("wrong: the right answer is billing"), and writes a
better instruction; the best of the instructions it writes is kept. This
is GEPA (Agrawal et al., 2025) with changes that suit a single function:
the model writing instructions sees what it tried that failed; a proposal
that copies an input instead of stating a rule is dropped; two
instructions that are right on different rows are combined; of equally
good instructions the shorter wins; and no row is run twice for the same
instruction. design/04-gepa.md in the repository gives the reasons.
Usage
gepa(
fn,
data,
expected = NULL,
metric = NULL,
feedback = NULL,
budget = 300L,
minibatch = 4L,
teacher = NULL,
selection = NULL,
seed = 0L
)Arguments
- fn
An AI function.
- data
Rows with the inputs and the right answers (in the columns named like the formula's outputs, or see
expected).- expected
Where the right answers are, as in
evaluate().- metric
A function
(row, prediction)returning a score, 1 meaning right, as inevaluate(). Default: exact match.- feedback
A function
(row, prediction, error)returning words about one answer (predictionis named like the outputs;errorisNULLunless the call failed). Default:"right","wrong: the right answer is ...", or the call's error.- budget
How many calls of
fnat most.- minibatch
Feedback rows per new instruction.
- teacher
The model that writes instructions (
"gpt-6-sol"). Default:fn's own.- selection
Rows to choose with, instead of half of
data.- seed
The random seed for splitting rows and picking parents.
Value
A copy of fn with the best instruction (fn itself, with its
search, when none beat the written one). ai_trials() gives the search.
Details
How it works. Half the rows (or selection) are for choosing and are
never shown to the model writing instructions; the other half are for
feedback. Every candidate instruction is scored on each choosing row. A
parent is picked among the candidates that are best on at least one row,
run on minibatch feedback rows, and teacher writes a new instruction
from its answers and their feedback. A new instruction that does better
on those rows is scored on the choosing rows and joins the candidates.
What it costs. budget calls of the function, plus one call of
teacher for each instruction it writes. Its score flatters: the
instruction kept is the best of many on the choosing rows, so measure it
on rows it never saw, with evaluate() (or resample it in tidymodels,
with set_engine("functai", method = "gepa")).
Examples
if (FALSE) { # \dontrun{
team <- ai(team ~ message, "Which team should answer this customer message?",
team = choice("shipping", "billing", "product", "account"))
better <- gepa(team, train, expected = category, teacher = "gpt-6-sol")
ai_instructions(better)
ai_trials(better)
evaluate(better, test, expected = category)
} # }