Skip to content

Make it cheaper

Compare models on your own data, count what each run costs, and cut tokens where it doesn't hurt.

import functai
functai.configure(lm="gpt-4.1-mini", temperature=0)   # the model behind every output on this page
from functai import ai, _ai

The best model for a job is the cheapest one that is right often enough, on your data. Benchmarks can't tell you which that is; twenty minutes with evaluate can.

The same question, several models

fn.using(lm=...) is the same function on another model. Run each on the same rows:

from typing import Literal
from dpyr import col, read

@ai
def team(message: str) -> Literal["shipping", "billing", "product", "account"]:
    """Which team should answer this customer message?

    House rules: anything wrong with the delivery itself, including an item
    that arrived broken, is shipping. Any request for money back is billing."""
    ...

tickets = functai.datasets.tickets()
results = []
for model in ["gpt-4.1-mini", "gpt-4.1-nano", "claude-haiku-4-5", "gemini-2.5-flash"]:
    ev = functai.evaluate(team.using(lm=model), tickets, expected="category", num_threads=8)
    summary = ev.summary.collect().to_dicts()[0]
    cost = ev.table.summarize(input_tokens=col.input_tokens.sum(), output_tokens=col.output_tokens.sum(),
                              seconds=col.seconds.mean()).collect().to_dicts()[0]
    results.append({"model": model, "right": ev.score, "low": summary["low"], "high": summary["high"], **cost})

models = read(results)
models
# dpyr dataframe · source: polars · showing 4 of 4 rows
┌──────────────────┬────────┬──────────┬──────────┬──────────────┬───────────────┬──────────┐
│ model            ┆ right  ┆ low      ┆ high     ┆ input_tokens ┆ output_tokens ┆ seconds  │
│ ---              ┆ ---    ┆ ---      ┆ ---      ┆ ---          ┆ ---           ┆ ---      │
│ str              ┆ f64    ┆ f64      ┆ f64      ┆ i64          ┆ i64           ┆ f64      │
╞══════════════════╪════════╪══════════╪══════════╪══════════════╪═══════════════╪══════════╡
│ gpt-4.1-mini     ┆ 0.9875 ┆ 0.932537 ┆ 0.99779  ┆ 7488         ┆ 720           ┆ 1.005348 │
│ gpt-4.1-nano     ┆ 0.875  ┆ 0.784972 ┆ 0.930664 ┆ 7488         ┆ 721           ┆ 1.060754 │
│ claude-haiku-4-5 ┆ 0.9375 ┆ 0.861899 ┆ 0.973011 ┆ 7741         ┆ 720           ┆ 0.573799 │
│ gemini-2.5-flash ┆ 0.9625 ┆ 0.895453 ┆ 0.987165 ┆ 7411         ┆ 400           ┆ 1.09125  │
└──────────────────┴────────┴──────────┴──────────┴──────────────┴───────────────┴──────────┘

Read right together with low and high: two models whose ranges overlap a lot may be equally good, and then the cheaper one wins. To be sure about two of them, compare their evaluations: pairing the rows settles it with fewer rows.

From tokens to money

functai counts tokens exactly; prices are your provider's, and they change, so bring your own (dollars per million tokens, input and output):

prices = {"gpt-4.1-mini": (0.40, 1.60), "gpt-4.1-nano": (0.10, 0.40),
          "claude-haiku-4-5": (1.00, 5.00), "gemini-2.5-flash": (0.30, 2.50)}   # check yours

for r in results:
    p_in, p_out = prices[r["model"]]
    dollars = (r["input_tokens"] * p_in + r["output_tokens"] * p_out) / 1e6
    r["per_1000_messages"] = round(dollars / len(tickets.collect()) * 1000, 3)
read(results).select(col.model, col.right, col.per_1000_messages)
# dpyr dataframe · source: polars · showing 4 of 4 rows
┌──────────────────┬────────┬───────────────────┐
│ model            ┆ right  ┆ per_1000_messages │
│ ---              ┆ ---    ┆ ---               │
│ str              ┆ f64    ┆ f64               │
╞══════════════════╪════════╪═══════════════════╡
│ gpt-4.1-mini     ┆ 0.9875 ┆ 0.052             │
│ gpt-4.1-nano     ┆ 0.875  ┆ 0.013             │
│ claude-haiku-4-5 ┆ 0.9375 ┆ 0.142             │
│ gemini-2.5-flash ┆ 0.9625 ┆ 0.04              │
└──────────────────┴────────┴───────────────────┘

(Those prices are illustrative, as listed by the providers in 2025; check your own before deciding.)

Where tokens go

  • Examples are sent with every call. Ten examples of 50 words each is 500 words on every row. Optimization that picks 4 good ones is often cheaper and better than 16.
  • Reasoning is paid output. An extra reasoning output, or module="cot", can help hard cases; measure whether it does before paying for it on every row.
  • Short instructions are fine. The rules that change answers are worth their words; politeness isn't.

Don't pay twice

While you work in a notebook, re-running a cell calls the model again. functai.configure(cache_replies=True) answers identical requests from memory instead (same model, same prompt, same input), so re-running an evaluation costs nothing; cache_replies="disk" keeps the replies across runs and processes, so a long run that was interrupted resumes where it stopped (Big tables). It's off by default because a cached answer hides how much a model's answers vary.

On tables, each distinct input is sent once per session anyway: 10,000 rows with 2,000 distinct messages cost 2,000 calls.

Very large volumes: your own small model

When a function runs millions of times, a small model can be trained to answer it: thousands of rows a second on one GPU, for nothing per call, with the unsure cases sent to a big model. See Bake it into a small model.