compare¶
compare(before, after)
Compare two evaluations of the same rows, row by row.
Pairing the rows detects a real change with far fewer examples than two separate scores would: a row both versions got right says nothing, a row only one got right says a lot.
Parameters¶
| Name | Type | Description | Default |
|---|---|---|---|
| before | Evaluation | Two evaluations of the same rows (typically two versions of a prompt, or two models). | required |
| after | Evaluation | Two evaluations of the same rows (typically two versions of a prompt, or two models). | required |
Returns¶
| Name | Type | Description |
|---|---|---|
| dpyr dataframe | One row per metric both have: before and after (the means), diff with its 95% interval low to high (a paired t interval), and how many rows got better, worse or stayed the same. When the interval includes 0, the change could be luck. |
See Also¶
evaluate: produces the evaluations.
Examples¶
import functai
from functai import *
from typing import Literal
rows = [
{"message": "The mug arrived in pieces.", "result": "shipping"},
{"message": "Box was crushed and the lamp inside is cracked.", "result": "shipping"},
{"message": "I want my money back for the toaster.", "result": "billing"},
{"message": "Please refund the blender, it stopped working.", "result": "billing"},
{"message": "Toaster burns one side of the bread.", "result": "product"},
]
@ai
def category(message: str) -> Literal["shipping", "billing", "product"]:
"""The support category of the message."""
...
@ai
def category_v2(message: str) -> Literal["shipping", "billing", "product"]:
"""The support category of the message. An item that arrived broken is
shipping; any request for money back is billing."""
...
compare(evaluate(category, rows, num_threads=5), evaluate(category_v2, rows, num_threads=5))
functai: no model chosen, so using gpt-4.1-mini (environment ($OPENAI_API_KEY)). Choose one with functai.configure(lm=...).
# dpyr dataframe · source: polars · showing 1 of 1 rows
┌─────────────┬────────┬───────┬──────┬───────────┬──────────┬────────┬───────┬──────┬─────┐
│ metric ┆ before ┆ after ┆ diff ┆ low ┆ high ┆ better ┆ worse ┆ same ┆ n │
│ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- │
│ str ┆ f64 ┆ f64 ┆ f64 ┆ f64 ┆ f64 ┆ i64 ┆ i64 ┆ i64 ┆ i64 │
╞═════════════╪════════╪═══════╪══════╪═══════════╪══════════╪════════╪═══════╪══════╪═════╡
│ exact_match ┆ 0.2 ┆ 0.8 ┆ 0.6 ┆ -0.079978 ┆ 1.279978 ┆ 3 ┆ 0 ┆ 2 ┆ 5 │
└─────────────┴────────┴───────┴──────┴───────────┴──────────┴────────┴───────┴──────┴─────┘