Measuring and improving

evaluate

e = evaluate(mood, rows)                         # rows: any table; inputs and answers by column name
e = evaluate(mood, rows; expected = :label)      # the answers are in another column
e.score, e.low, e.high                           # the mean and its 95% interval
DataFrame(e)                                     # every row: its columns, pred_<output>, its score, error, call

Each row's inputs are its columns named like the function's inputs. A row whose call fails scores 0 and keeps its error. The calls carry caller.evaluation = e.run in the call log, so they are never mistaken for real use.

The default metric, exact_match, compares text ignoring case and runs of white space (Unicode case folding, as every FunctAI language does it), numbers by value, and an @enum or Symbol answer with text by its name:

exact_match(Dict("result" => "Billing "), Dict("result" => "billing"))
OrderedCollections.OrderedDict{String, Float64} with 1 entry:
  "exact_match" => 1.0

Any other metric is a function of the row and the outputs, each a NamedTuple:

evaluate(solve, rows; metric = (row, out) -> abs(out.result - row.answer) < 0.01)
evaluate(triage, rows; metric = Dict("summary_ok" => (row, out) -> length(out.summary) < 120,
                                     "minutes_ok" => (row, out) -> out.minutes == row.minutes))

evaluate takes any function: a plain Julia one is called with each row, so baselines are measured the same way:

always_billing(row) = "billing"
rows = [(message = "Charged twice", result = "billing"), (message = "Late parcel", result = "shipping")]
evaluate(always_billing, rows)
Evaluation of always_billing on 2 rows
  exact_match  0.50  (95% range 0.09 to 0.91)
  every row: DataFrame(e)

The interval

Right-or-wrong scores get Wilson's interval; any other score Student's t. Both are computed exactly as in every FunctAI language, so a score measured in Julia compares with one measured in R:

score_interval(vcat(ones(72), zeros(8)))
(mean = 0.9, low = 0.8148931111226091, high = 0.9484523846972116)

Two versions, row by row

On the same rows, only the rows where two versions disagree say which is better. compare counts them and gives the mean difference with a paired interval:

compare(evaluate(team, rows), evaluate(team_rules, rows))

Improving

Each optimizer returns an improved copy; the function you pass is unchanged. Only the instruction and the worked examples change, and so does the version.

What it doesCalls
with_demos, with_instructionsset them by handnone
labeled_few_shotk rows with known answers become worked examplesnone
bootstrap_few_shotruns the function (or a teacher model) on rows; the runs the metric accepts become worked examples, reasoning and tool calls includedone per row tried
random_searchseveral sets of examples, each scored on valset; the best winsmany
gepaa stronger model (teacher) reads the function's answers on a few rows, with feedback in words, and rewrites the instruction; candidates are kept in a Pareto pool scored on selection, combined, and the best (of equals, the shortest) winsup to budget
instruction_searcha stronger model (prompt_lm) proposes instructions; each is tried on minibatches of valset; the best winsmany
better = bootstrap_few_shot(refund, train; teacher = "gpt-6-sol", max_bootstrapped = 4)
rewritten, trials = gepa(refund, train; selection = dev, teacher = "gpt-6-sol", budget = 300)
DataFrame(trials)                                # every instruction tried, and why it was kept or dropped

Keep three piles of rows: one to learn from, one to choose on, one to test once. Choosing the best of several versions on the same rows flatters the winner, and a search is itself random: measure its result on rows it never saw. AIModel(method = :gepa) runs the search inside fit!, so MLJ's cross-validation measures the search itself.

Datasets

Three small labelled tables to learn with, the same as Python's and R's: FunctAI.tickets, FunctAI.field_notes, FunctAI.refunds. Each docstring has the rules its labels follow.

first(DataFrame(FunctAI.tickets()), 3)
3×5 DataFrame
Rowidmessagechannelcategoryorder_id
Int64StringStringStringString?
11Hi, my order A-1042 still hasn't arrived and it's been three weeks.emailshippingA-1042
22The mug arrived in pieces.chatshippingmissing
33I was charged twice for order B-2210, please fix this.emailbillingB-2210