Skip to content

rated

rated(program, *, folder=None, by=None, since=None, any_file=False)

The calls people rated, as rows with known answers.

Each row is a call someone judged: its inputs under their names, and the right answer under the output's name (the answer it gave, when it was marked right; the correction, when it was marked wrong and corrected). Hand it to evaluate or .opt as it is.

A call marked wrong without the right answer is left out: it says what the answer isn't, not what it is. So are calls made when the program had other inputs or outputs (another signature), and calls logged without their values; a warning says how many.

Rated calls are the ones people chose to look at, not a fair sample: a score on them says how the program does on those. To measure how often it is right, rate a random draw (rate(..., sample=...)) and keep the rows of that draw (col.sample == "...").

Parameters

Name Type Description Default
program AI function, module or name Whose calls. Given the function itself, only calls with its current signature are used (calls from earlier versions of the prompt are kept: their inputs and outputs are the same). required
folder str or path The log folder (default as for calls). None
by str Only this person's ratings. Default: everyone's; when people disagree, the latest rating is used and disputed is true. None
since (date, datetime, timedelta or text) Only calls and ratings from then on. None
any_file bool A program defined in a notebook or a script is known by its file too, so two notebooks' summarize are two programs. True takes its calls from any file (a notebook that was moved or renamed). False

Returns

Name Type Description
dpyr dataframe The inputs, the right answer (named like the output), then rating (right or wrong), rated_by, origin (review or edit), disputed, sample, version and call.

See Also

  • rate: judge a call.
  • calls: every logged call.
  • evaluate: score a program on these rows.

Examples

import functai
from functai import *
import tempfile
from typing import Literal

@ai
def team(message: str) -> Literal["shipping", "billing", "product"]:
    """Which team should answer this customer message?"""
    ...

with functai.configure(log_calls=tempfile.mkdtemp()):
    a = team.predict("My parcel never came.")
    b = team.predict("The chair arrived with a snapped leg.")
    functai.rate(a, "right")
    functai.rate(b, "wrong", answer="shipping")    # broken on the way: the carrier's fault
    rows = functai.rated(team)
rows.select("message", "result", "rating")
functai: no model chosen, so using gpt-4.1-mini (environment ($OPENAI_API_KEY)). Choose one with functai.configure(lm=...).
# dpyr dataframe · source: polars · showing 2 of 2 rows
┌───────────────────────────────────────┬──────────┬────────┐
│ message                               ┆ result   ┆ rating │
│ ---                                   ┆ ---      ┆ ---    │
│ str                                   ┆ str      ┆ str    │
╞═══════════════════════════════════════╪══════════╪════════╡
│ My parcel never came.                 ┆ shipping ┆ right  │
│ The chair arrived with a snapped leg. ┆ shipping ┆ wrong  │
└───────────────────────────────────────┴──────────┴────────┘