Skip to content

7. A model you own

A bank gets thousands of questions a day in 77 kinds. By the end you will have trained a small model that answers the same AI function on your own computer, for nothing per call and thousands of times faster; measured it against the language model it replaces; and set it to hand only the questions it's unsure about to the language model.

Can you skip this one? If you can answer these, jump to tutorial 8. The answers are at the bottom.

  1. What changes in an AI function when you bake it, and what doesn't?
  2. You have no labelled rows. Where do a small model's labels come from, and what limits how good it can get?
  3. How do you keep a cheap model's accuracy high without paying for a big model on every question?

What you need

This is the one tutorial that trains weights. It needs functai's bake extra (PyTorch and Hugging Face transformers, a few gigabytes) and, in practice, a GPU: the training below takes about a minute on one RTX 3090 and much longer on a laptop's CPU. It downloads a 17-million-parameter model (about 70 MB) and a public dataset. About 25 cents of model calls.

pip install "functai[data,bake]"
import tempfile
from typing import Literal

from dpyr import col, n, read

import functai
from functai import ai

log_folder = tempfile.mkdtemp()
functai.configure(lm="gpt-6-luna", log_calls=log_folder)
configure(lm='gpt-6-luna', log_calls='/tmp/tmpnlx1lqco')

The questions

banking77 is a public dataset of 13,000 real questions to a bank, each labelled with one of 77 intents by people. It's a fair stand-in for any high-volume classification job: support tickets, forms, logs.

url = "https://raw.githubusercontent.com/PolyAI-LDN/task-specific-datasets/master/banking_data/"
train = read(url + "train.csv")
test = read(url + "test.csv")
train
# dpyr dataframe · source: polars · showing 10 of ? rows
┌──────────────────────────────────────────────────────────────┬──────────────┐
│ text                                                         ┆ category     │
│ ---                                                          ┆ ---          │
│ str                                                          ┆ str          │
╞══════════════════════════════════════════════════════════════╪══════════════╡
│ I am still waiting on my card?                               ┆ card_arrival │
│ What can I do if my card still hasn't arrived after 2 weeks? ┆ card_arrival │
│ I have been waiting over a week. Is the card still coming?   ┆ card_arrival │
│ Can I track my card while it is in the process of delivery?  ┆ card_arrival │
│ How do I know if I will get my card, or if it is lost?       ┆ card_arrival │
│ When did you send me my new card?                            ┆ card_arrival │
│ Do you have info about the card on delivery?                 ┆ card_arrival │
│ What do I do if I still have not received my new card?       ┆ card_arrival │
│ Does the package with my card have tracking?                 ┆ card_arrival │
│ I ordered my card but it still isn't here                    ┆ card_arrival │
└──────────────────────────────────────────────────────────────┴──────────────┘
intents = sorted(set(train.pull(col.category)))
len(intents), intents[:8]
(77, ['Refund_not_showing_up', 'activate_my_card', 'age_limit', 'apple_pay_or_google_pay', 'atm_support', 'automatic_top_up', 'balance_not_updated_after_bank_transfer', 'balance_not_updated_after_cheque_or_cash_deposit'])

Seventy-seven answers is too many to type, so build the Literal from the list. The function is as short as ever:

Intent = Literal[tuple(intents)]

@ai
def intent(text: str) -> Intent:
    """What the bank's customer wants."""
    ...

We'll measure everything on the same 300 test questions (with the answers in a column named like the output, result), and keep the rest of the test set for the training reports.

labelled_test = test.rename(result=col.category)
measure_on = labelled_test.slice_sample(n=300, seed=7)

The language model, measured

ev_luna = functai.evaluate(intent, measure_on, num_threads=16)
ev_luna
Evaluation(intent, 300 examples: exact_match 0.81 [0.77, 0.85])
ev_luna.table.summarize(seconds_each=col.seconds.median(),
                        dollars_per_1000=(1000 * (col.input_tokens * 0.10 + (col.total_tokens - col.input_tokens) * 0.50) / 1e6).mean())
# dpyr dataframe · source: polars · showing 1 of 1 rows
┌──────────────┬──────────────────┐
│ seconds_each ┆ dollars_per_1000 │
│ ---          ┆ ---              │
│ f64          ┆ f64              │
╞══════════════╪══════════════════╡
│ 1.628349     ┆ 0.075085         │
└──────────────┴──────────────────┘

Good, and cheap per question. But a bank that gets a million questions a month pays for a million calls, waits a second for each, and sends every customer's words to another company. For a job this narrow, there is another way.

Bake it

Baking trains a small model to answer an AI function, then runs the same function on those weights. The function doesn't change (same name, same types, same answers); what executes it does. The student here is Ettin, a 17-million-parameter encoder: it reads the question and gives a probability for each of the 77 answers.

First, with the people's labels, all 10,000 training questions:

people = intent.bake(train.rename(result=col.category), test=labelled_test.slice_head(n=1000), log=False)
print(people.report)
Baked intent: jhu-clsp/ettin-encoder-17m (16.9M parameters)
  trained on 9,003 rows (the data's labels), validated on 1,000, tested on 1,000 labeled rows
  training: 6 passes (best 6), 279 s on cpu (fp32), inputs up to 56 tokens

  on the test rows            student                 
  accuracy                    91.1% (89.2%–92.7%)     
  top-3                       97.1%
  calibration error (ECE)     0.020 (was 0.051; temperature 1.66)

  answering only when sure:  most confident share → accuracy (confidence at the cut)
      50% → 99.6%  (≥ 0.99)
      80% → 98.1%  (≥ 0.89)
      90% → 96.3%  (≥ 0.68)
     100% → 91.1%  (≥ 0.12)
    for 95% accuracy: escalate_below=0.59 keeps 93% of rows

  speed on cpu: 2,434 rows/s batched (tokenizing included), 3.6 ms for one row

Read the report top to bottom: what it was trained on and how long it took, its accuracy on the test rows with an interval, its calibration error (how far its confidences are from the truth; lower is better, and a fitted temperature corrects it), the accuracy you get if it only answers when it's sure, and its speed.

Now run the same function on the baked weights. using(lm=...) takes a baked model like any model name:

fast = intent.using(lm=people)

p = fast.predict("my card still hasn't arrived after two weeks")
p.result, round(p.confidence, 3)
('card_arrival', 0.989)

On the same 300 questions:

ev_fast = functai.evaluate(fast, measure_on, num_threads=16)
ev_fast
Evaluation(intent, 300 examples: exact_match 0.89 [0.85, 0.92])
functai.compare(ev_luna, ev_fast)
# dpyr dataframe · source: polars · showing 1 of 1 rows
┌─────────────┬──────────┬──────────┬──────────┬──────────┬──────────┬────────┬───────┬──────┬─────┐
│ metric      ┆ before   ┆ after    ┆ diff     ┆ low      ┆ high     ┆ better ┆ worse ┆ same ┆ n   │
│ ---         ┆ ---      ┆ ---      ┆ ---      ┆ ---      ┆ ---      ┆ ---    ┆ ---   ┆ ---  ┆ --- │
│ str         ┆ f64      ┆ f64      ┆ f64      ┆ f64      ┆ f64      ┆ i64    ┆ i64   ┆ i64  ┆ i64 │
╞═════════════╪══════════╪══════════╪══════════╪══════════╪══════════╪════════╪═══════╪══════╪═════╡
│ exact_match ┆ 0.813333 ┆ 0.886667 ┆ 0.073333 ┆ 0.026692 ┆ 0.119975 ┆ 37     ┆ 15    ┆ 248  ┆ 300 │
└─────────────┴──────────┴──────────┴──────────┴──────────┴──────────┴────────┴───────┴──────┴─────┘

The paired comparison says whether your own model is clearly worse, clearly better, or indistinguishable from the language model on these questions. On banking77 the small model trained on people's labels comes out ahead, and that isn't a fluke: 77 intents drawn where this bank's labellers drew them are exactly what a language model has to guess, and exactly what labels teach. And it costs nothing per call, runs on your hardware, and answers in milliseconds.

No labels? A teacher

Most jobs don't start with 10,000 labelled rows. Then the language model can be the teacher: it labels unlabelled questions once, and the student learns from its labels. Here, 2,000 training questions with the labels thrown away:

unlabelled = train.select(col.text).slice_sample(n=2000, seed=1)

taught = intent.bake(unlabelled, teacher="gpt-6-luna", test=labelled_test.slice_head(n=500),
                     compare_teacher=True, prices={"teacher": (0.10, 0.50)}, log=False)
print(taught.report)
Baked intent: jhu-clsp/ettin-encoder-17m (16.9M parameters)
  trained on 1,800 rows (teacher (hard)), validated on 200, tested on 500 labeled rows
  training: 6 passes (best 4), 58 s on cpu (fp32), inputs up to 56 tokens

  on the test rows            student                 teacher (gpt-6-luna)
  accuracy                    71.0% (66.9%–74.8%)     86.2%
  top-3                       85.6%
  calibration error (ECE)     0.057 (was 0.130; temperature 1.48)
  agrees with the teacher     74.0%

  answering only when sure:  most confident share → accuracy (confidence at the cut)
      50% → 86.8%  (≥ 0.79)
      80% → 79.0%  (≥ 0.46)
      90% → 76.0%  (≥ 0.33)
     100% → 71.0%  (≥ 0.11)
    for 95% accuracy: escalate_below=0.98 keeps 9% of rows

  speed on cpu: 1,644 rows/s batched (tokenizing included), 3.7 ms for one row
  teacher labels: 2,000 rows from gpt-6-luna in 234 s (507 tokens a row, $0.15)
  break-even in time against the teacher: after 2,514 rows

  note: result: the student (71.0%) is below its teacher (86.2%); trained on teacher labels, it can at best match it. Human labels lifted the same kind of student from 77% to 91.5% on banking77: label more rows by hand, or use a stronger teacher

With compare_teacher=True (it costs a teacher pass over the test rows, so it is off unless asked), the report shows the teacher next to the student, on the same test rows, and what the labels cost. It also says the thing to remember: a student trained on a teacher's labels can at best match the teacher, and usually lands a little below it. People's labels, when you have them, are worth more than any teacher's.

Small model first, language model when unsure

The student's confidence is calibrated, so you can use it the way tutorial 6 used Jev's: answer when sure, and send the rest to the language model. The report tells you where to cut for a target accuracy:

cut = people.report.threshold(0.95)
cut
{'threshold': 0.5874377718045757, 'share': 0.93, 'accuracy': 0.9505376344086022}
careful = intent.using(lm=people, escalate_to="gpt-6-luna", escalate_below=cut["threshold"])

ev_careful = functai.evaluate(careful, measure_on, num_threads=16)
escalated = sum(bool(p and p.escalated) for p in ev_careful.predictions)
ev_careful, f"{escalated} of {len(ev_careful)} asked gpt-6-luna"
(Evaluation(intent, 300 examples: exact_match 0.92 [0.88, 0.94]), '27 of 300 asked gpt-6-luna')
read([{"setup": name, **ev.summary.collect().to_dicts()[0]} for name, ev in [
    ("gpt-6-luna on everything", ev_luna),
    ("your model alone", ev_fast),
    ("your model, gpt-6-luna when unsure", ev_careful),
]]).select(col.setup, col.mean, col.low, col.high)
# dpyr dataframe · source: polars · showing 3 of 3 rows
┌────────────────────────────────────┬──────────┬──────────┬──────────┐
│ setup                              ┆ mean     ┆ low      ┆ high     │
│ ---                                ┆ ---      ┆ ---      ┆ ---      │
│ str                                ┆ f64      ┆ f64      ┆ f64      │
╞════════════════════════════════════╪══════════╪══════════╪══════════╡
│ gpt-6-luna on everything           ┆ 0.813333 ┆ 0.765381 ┆ 0.853363 │
│ your model alone                   ┆ 0.886667 ┆ 0.845801 ┆ 0.917755 │
│ your model, gpt-6-luna when unsure ┆ 0.916667 ┆ 0.879878 ┆ 0.942919 │
└────────────────────────────────────┴──────────┴──────────┴──────────┘

Most questions are answered by your model, for free; the language model sees only the ones your model was unsure of. Compare the accuracy and the number of paid calls: that's the trade you tune with escalate_below.

Keep it

A baked model is a folder: the weights, the tokenizer, and what it was trained to answer. Load it back anywhere functai runs:

import os
folder = os.path.join(tempfile.mkdtemp(), "intent-model")
people.save(folder)

from functai.bake import load
reloaded = load(folder)
intent.using(lm=reloaded)("How do I top up with Apple Pay?")
'apple_pay_or_google_pay'

A function whose inputs, outputs or answers change is refused by the baked model it no longer matches, so a stale model can't silently answer the wrong question. functai.save() of a program copies its baked weights along (tutorial 8).

What it cost

functai.calls(folder=log_folder).summarize(
    calls=n(), dollars=((col.input_tokens * 0.10 + (col.total_tokens - col.input_tokens) * 0.50) / 1e6).sum())
# dpyr dataframe · source: polars · showing 1 of 1 rows
┌───────┬───────────┐
│ calls ┆ dollars   │
│ ---   ┆ ---       │
│ i64   ┆ f64       │
╞═══════╪═══════════╡
│ 3402  ┆ 0.2148021 │
└───────┴───────────┘

The training itself cost only electricity: a minute of GPU.

Your turn

  1. Bake with teacher="jev-latest" instead. Jev gives a probability per answer, which the student learns from as a soft label. Is the student better than with gpt-6-luna's labels? Cheaper?
  2. Try escalate_below at 0.5, 0.8 and 0.95 of people's confidence. Plot paid calls against accuracy.
  3. Bake with only 500 of the people's labels. How much accuracy does the data buy? (That curve tells you how much labelling is worth.)

What you learned

  • fn.bake(rows) trains a small model to answer an AI function; fn.using(lm=baked) runs the same function on it, on your hardware, for nothing per call.
  • The report measures the student on held-out rows: accuracy with an interval, calibration, accuracy when answering only when sure, speed.
  • Without labels, a teacher= model labels the rows once; the student can at best match its teacher. People's labels are worth more.
  • A calibrated student can escalate: escalate_to= a language model, below a confidence the report suggests.
  • A baked model is a folder: save(), functai.bake.load().

Answers to the check at the top. (1) What executes it: the name, types, docstring and answers stay; the weights answering change. (2) From a teacher (a language model, or Jev) that labels them once; the student can at best match the teacher. (3) Answer with the small model when it's sure, and escalate the rest (escalate_to=, escalate_below=).

Next: 8. Living with it: tools, the call log, people's corrections, versions, and saving a whole program.