Skip to contents

A model reads the function's answers on rows with known answers, with feedback in words ("wrong: the right answer is billing"), and writes a better instruction; the best of the instructions it writes is kept. This is GEPA (Agrawal et al., 2025) with changes that suit a single function: the model writing instructions sees what it tried that failed; a proposal that copies an input instead of stating a rule is dropped; two instructions that are right on different rows are combined; of equally good instructions the shorter wins; and no row is run twice for the same instruction. design/04-gepa.md in the repository gives the reasons.

Usage

gepa(
  fn,
  data,
  expected = NULL,
  metric = NULL,
  feedback = NULL,
  budget = 300L,
  minibatch = 4L,
  teacher = NULL,
  selection = NULL,
  seed = 0L
)

Arguments

fn

An AI function.

data

Rows with the inputs and the right answers (in the columns named like the formula's outputs, or see expected).

expected

Where the right answers are, as in evaluate().

metric

A function (row, prediction) returning a score, 1 meaning right, as in evaluate(). Default: exact match.

feedback

A function (row, prediction, error) returning words about one answer (prediction is named like the outputs; error is NULL unless the call failed). Default: "right", "wrong: the right answer is ...", or the call's error.

budget

How many calls of fn at most.

minibatch

Feedback rows per new instruction.

teacher

The model that writes instructions ("gpt-6-sol"). Default: fn's own.

selection

Rows to choose with, instead of half of data.

seed

The random seed for splitting rows and picking parents.

Value

A copy of fn with the best instruction (fn itself, with its search, when none beat the written one). ai_trials() gives the search.

Details

How it works. Half the rows (or selection) are for choosing and are never shown to the model writing instructions; the other half are for feedback. Every candidate instruction is scored on each choosing row. A parent is picked among the candidates that are best on at least one row, run on minibatch feedback rows, and teacher writes a new instruction from its answers and their feedback. A new instruction that does better on those rows is scored on the choosing rows and joins the candidates.

What it costs. budget calls of the function, plus one call of teacher for each instruction it writes. Its score flatters: the instruction kept is the best of many on the choosing rows, so measure it on rows it never saw, with evaluate() (or resample it in tidymodels, with set_engine("functai", method = "gepa")).

Examples

if (FALSE) { # \dontrun{
team <- ai(team ~ message, "Which team should answer this customer message?",
  team = choice("shipping", "billing", "product", "account"))
better <- gepa(team, train, expected = category, teacher = "gpt-6-sol")
ai_instructions(better)
ai_trials(better)
evaluate(better, test, expected = category)
} # }