Measuring and improving
evaluate
e = evaluate(mood, rows) # rows: any table; inputs and answers by column name
e = evaluate(mood, rows; expected = :label) # the answers are in another column
e.score, e.low, e.high # the mean and its 95% interval
DataFrame(e) # every row: its columns, pred_<output>, its score, error, callEach row's inputs are its columns named like the function's inputs. A row whose call fails scores 0 and keeps its error. The calls carry caller.evaluation = e.run in the call log, so they are never mistaken for real use.
The default metric, exact_match, compares text ignoring case and runs of white space (Unicode case folding, as every FunctAI language does it), numbers by value, and an @enum or Symbol answer with text by its name:
exact_match(Dict("result" => "Billing "), Dict("result" => "billing"))OrderedCollections.OrderedDict{String, Float64} with 1 entry:
"exact_match" => 1.0Any other metric is a function of the row and the outputs, each a NamedTuple:
evaluate(solve, rows; metric = (row, out) -> abs(out.result - row.answer) < 0.01)
evaluate(triage, rows; metric = Dict("summary_ok" => (row, out) -> length(out.summary) < 120,
"minutes_ok" => (row, out) -> out.minutes == row.minutes))evaluate takes any function: a plain Julia one is called with each row, so baselines are measured the same way:
always_billing(row) = "billing"
rows = [(message = "Charged twice", result = "billing"), (message = "Late parcel", result = "shipping")]
evaluate(always_billing, rows)Evaluation of always_billing on 2 rows
exact_match 0.50 (95% range 0.09 to 0.91)
every row: DataFrame(e)The interval
Right-or-wrong scores get Wilson's interval; any other score Student's t. Both are computed exactly as in every FunctAI language, so a score measured in Julia compares with one measured in R:
score_interval(vcat(ones(72), zeros(8)))(mean = 0.9, low = 0.8148931111226091, high = 0.9484523846972116)Two versions, row by row
On the same rows, only the rows where two versions disagree say which is better. compare counts them and gives the mean difference with a paired interval:
compare(evaluate(team, rows), evaluate(team_rules, rows))Improving
Each optimizer returns an improved copy; the function you pass is unchanged. Only the instruction and the worked examples change, and so does the version.
| What it does | Calls | |
|---|---|---|
with_demos, with_instructions | set them by hand | none |
labeled_few_shot | k rows with known answers become worked examples | none |
bootstrap_few_shot | runs the function (or a teacher model) on rows; the runs the metric accepts become worked examples, reasoning and tool calls included | one per row tried |
random_search | several sets of examples, each scored on valset; the best wins | many |
gepa | a stronger model (teacher) reads the function's answers on a few rows, with feedback in words, and rewrites the instruction; candidates are kept in a Pareto pool scored on selection, combined, and the best (of equals, the shortest) wins | up to budget |
instruction_search | a stronger model (prompt_lm) proposes instructions; each is tried on minibatches of valset; the best wins | many |
better = bootstrap_few_shot(refund, train; teacher = "gpt-6-sol", max_bootstrapped = 4)
rewritten, trials = gepa(refund, train; selection = dev, teacher = "gpt-6-sol", budget = 300)
DataFrame(trials) # every instruction tried, and why it was kept or droppedKeep three piles of rows: one to learn from, one to choose on, one to test once. Choosing the best of several versions on the same rows flatters the winner, and a search is itself random: measure its result on rows it never saw. AIModel(method = :gepa) runs the search inside fit!, so MLJ's cross-validation measures the search itself.
Datasets
Three small labelled tables to learn with, the same as Python's and R's: FunctAI.tickets, FunctAI.field_notes, FunctAI.refunds. Each docstring has the rules its labels follow.
first(DataFrame(FunctAI.tickets()), 3)| Row | id | message | channel | category | order_id |
|---|---|---|---|---|---|
| Int64 | String | String | String | String? | |
| 1 | 1 | Hi, my order A-1042 still hasn't arrived and it's been three weeks. | shipping | A-1042 | |
| 2 | 2 | The mug arrived in pieces. | chat | shipping | missing |
| 3 | 3 | I was charged twice for order B-2210, please fix this. | billing | B-2210 |