Skip to content

Make it better

Rules in the docstring, examples, optimizers, bigger models: in order of cost, each one measured.

import functai
functai.configure(lm="gpt-4.1-mini", temperature=0)   # the model behind every output on this page
from functai import ai, _ai

When a function isn't right often enough, there's a ladder of things to try, from free to expensive. Climb it in order, and measure each step on rows the function didn't learn from.

step costs good when
1. say the rule in the docstring nothing you can put the rule into words
2. a tighter type nothing answers drift out of the allowed set, or have the wrong shape
3. examples, by hand a few tokens per call the rule is easier to show than to say
4. examples chosen from your data (functai.bootstrap_few_shot) one run over the data you have labelled rows
5. an optimizer that also rewrites the instruction many runs steps 1 to 4 plateau
6. a bigger model every call nothing else moved it

Two sets of rows

Learn from some rows, judge on the others. Here, the first 40 tickets to learn from and the last 40 to judge:

from typing import Literal
from dpyr import col

tickets = functai.datasets.tickets()
learn = tickets.filter(col.id <= 40)
judge_on = tickets.filter(col.id > 40)

@ai
def team(message: str) -> Literal["shipping", "billing", "product", "account"]:
    """Which team should answer this customer message?"""
    ...

start = functai.evaluate(team, judge_on, expected="category", num_threads=8)
start
Evaluation(team, 40 examples: exact_match 0.90 [0.77, 0.96])

1. Say the rule

Read the misses first. If you can say what they have in common, say it in the docstring: it's free, exact, and anyone can read it later. (This is the step Get started takes.)

2. A tighter type

A Literal or an Enum restricts the answer to a set; a dataclass gives it named parts; int | None allows "no answer" instead of a guess. The type is part of the prompt, and the reply is checked against it. See Types.

3. Examples, by hand

A few worked examples, shown before every question:

@ai(examples=[
    ("The vase came smashed.", "shipping"),
    ("Money back please, the chair wobbles.", "billing"),
])
def team_ex(message: str) -> Literal["shipping", "billing", "product", "account"]:
    """Which team should answer this customer message?"""
    ...

functai.compare(start, functai.evaluate(team_ex, judge_on, expected="category", num_threads=8))
# dpyr dataframe · source: polars · showing 1 of 1 rows
┌─────────────┬────────┬───────┬────────┬───────────┬──────────┬────────┬───────┬──────┬─────┐
│ metric      ┆ before ┆ after ┆ diff   ┆ low       ┆ high     ┆ better ┆ worse ┆ same ┆ n   │
│ ---         ┆ ---    ┆ ---   ┆ ---    ┆ ---       ┆ ---      ┆ ---    ┆ ---   ┆ ---  ┆ --- │
│ str         ┆ f64    ┆ f64   ┆ f64    ┆ f64       ┆ f64      ┆ i64    ┆ i64   ┆ i64  ┆ i64 │
╞═════════════╪════════╪═══════╪════════╪═══════════╪══════════╪════════╪═══════╪══════╪═════╡
│ exact_match ┆ 0.9    ┆ 0.875 ┆ -0.025 ┆ -0.139223 ┆ 0.089223 ┆ 2      ┆ 3     ┆ 35   ┆ 40  │
└─────────────┴────────┴───────┴────────┴───────────┴──────────┴────────┴───────┴──────┴─────┘

Pairs of (input, answer), or rows like {"message": ..., "result": ...}.

4. Examples chosen from your data

functai.bootstrap_few_shot runs the function on your labelled rows and keeps up to 4 runs that got the right answer as worked examples (with their reasoning and tool calls, if any), then adds labelled rows up to 16 examples. It returns an improved copy: team itself is unchanged, so you can measure the two side by side. Your code, types and prompt format are never touched.

taught = functai.bootstrap_few_shot(team, learn, expected="category")
taught.state()
instruction: (written from the code)
examples: 16
  1. message="Hi, my order A-1042 still hasn't arrived and it's been three weeks."  →  result='shipping'
  2. message='The mug arrived in pieces.'  →  result='shipping'
  3. message='I was charged twice for order B-2210, please fix this.'  →  result='billing'
  4. message='How do I change the email on my account?'  →  result='account'
  5. message="When will order B-2417 ship? It says 'processing' for a week."  →  result='shipping'
  6. message='Can you merge my two accounts? I signed up twice by mistake.'  →  result='account'
  7. message='Two of the six plates were broken on arrival, order D-4120.'  →  result='shipping'
  8. message='I want a refund for the chair, it wobbles no matter what I do.'  →  result='billing'
  9. message="My coupon code SPRING10 didn't apply at checkout."  →  result='billing'
  10. message='The glass carafe was shattered when I opened the package.'  →  result='shipping'
  11. message='The non-stick coating is peeling off my frying pan.'  →  result='product'
  12. message='The vase arrived with a big crack down the side.'  →  result='shipping'
  13. message="Order c3319 was delivered to my neighbour's address instead of mine."  →  result='shipping'
  14. message='i cant log in it says my account is locked??'  →  result='account'
  15. message='Is the cutting board safe to use for raw meat?'  →  result='product'
  16. message='Does the stand mixer come with a dough hook?'  →  result='product'
optimized = functai.compare(start, functai.evaluate(taught, judge_on, expected="category", num_threads=8))
optimized
# dpyr dataframe · source: polars · showing 1 of 1 rows
┌─────────────┬────────┬───────┬───────┬───────────┬──────────┬────────┬───────┬──────┬─────┐
│ metric      ┆ before ┆ after ┆ diff  ┆ low       ┆ high     ┆ better ┆ worse ┆ same ┆ n   │
│ ---         ┆ ---    ┆ ---   ┆ ---   ┆ ---       ┆ ---      ┆ ---    ┆ ---   ┆ ---  ┆ --- │
│ str         ┆ f64    ┆ f64   ┆ f64   ┆ f64       ┆ f64      ┆ i64    ┆ i64   ┆ i64  ┆ i64 │
╞═════════════╪════════╪═══════╪═══════╪═══════════╪══════════╪════════╪═══════╪══════╪═════╡
│ exact_match ┆ 0.9    ┆ 0.975 ┆ 0.075 ┆ -0.036904 ┆ 0.186904 ┆ 4      ┆ 1     ┆ 35   ┆ 40  │
└─────────────┴────────┴───────┴───────┴───────────┴──────────┴────────┴───────┴──────┴─────┘

On these 40 judging rows: +5 points (3 rows better, 1 worse), somewhere between −5 and +15. That range includes 0, so with 40 rows this can't be told apart from luck; more judging rows would settle it. Measuring is what keeps you from shipping a change that only looked better.

team is still the function you started with; taught.optimization_runs() says how the copy was made. To keep the result, save the function (or just its examples: taught.save("team.json")).

5. Instructions, rewritten and tried

InstructionSearch asks a model to propose instructions from your code and a few rows, tries each (with sets of examples) on small batches, and keeps the one that scores best on the judging rows. It costs many runs; use it when steps 1 to 4 have stopped helping.

from functai import InstructionSearch

search = InstructionSearch(num_candidates=6, num_trials=12)
searched = team.opt(learn, valset=judge_on, expected="category", optimizer=search)
searched.trials      # every try: which instruction, which examples, its score

The translator example runs it end to end, with a model as the judge.

GEPA rewrites the instruction from the function's own mistakes instead: a teacher model reads its answers on a few rows with feedback in words ("wrong: the right answer is billing"), writes a better instruction, and the best of what it writes is kept, chosen on rows it is never shown. It is at its best when a small, cheap model runs the function and a large one writes its instruction once:

small = team.using(lm="gpt-5.4-nano")                  # the model that will run it
better = functai.gepa(small, learn, expected="category",
                      teacher="gpt-6-sol", budget=300)  # the model that writes its instruction
better.instructions
better.trials        # every instruction tried: its parent, how it was made, its score, its length

Its score on the rows it chose with flatters (it is the best of many there): measure it on rows it never saw. How it differs from the paper's GEPA, and why, is in design/04-gepa.md.

6. A bigger model, or a teacher

team.using(lm="gpt-4.1")                            # a bigger model for every call
functai.bootstrap_few_shot(team, learn, expected="category",
                           teacher="gpt-4.1")       # or: a big model writes the examples once,
                                                    # and the small one uses them from then on

The second line is often the better deal: you pay for the big model once, during optimization, not on every call. Make it cheaper shows how to check.

The optimizers

By name, the common cases: functai.labeled_few_shot(fn, rows, k=8), functai.bootstrap_few_shot(fn, rows, teacher=...), functai.gepa(fn, rows, teacher=...). Each returns an improved copy. For the others, or their every option, fn.opt(rows, optimizer=...) (it uses BootstrapFewShot unless told otherwise) also returns a copy:

optimizer= what it does
LabeledFewShot(k=16) your labelled rows, as examples
BootstrapFewShot(...) (default) runs the function; runs that were right become examples, reasoning and tool calls included
BootstrapFewShotWithRandomSearch(num_candidate_programs=8) several sets of examples, keeps the best on valset
InstructionSearch(num_candidates=6, num_trials=12) proposed instructions × sets of examples, keeps the best on valset
GEPA(budget=300, teacher=...) a teacher rewrites the instruction from the mistakes, with feedback in words; keeps the best on rows it never shows

Every option is in the reference.