Skip to contents

Runs a model on every row of data and scores each answer against the right one, with a 95% interval. The model is an AI function, a fitted parsnip model, a fitted workflow, or anything with a predict() method returning .pred_class or .pred: the same score, the same interval, so an AI function and a classical model compare on equal terms. A row whose call failed scores 0. An AI function's calls are logged with caller.evaluation set, so they are not mistaken for use.

Usage

evaluate(fn, data, expected = NULL, metric = NULL, ...)

Arguments

fn

A model: an AI function, a parsnip fit, a workflow, ...

data

A data frame: the model's inputs, and the right answers.

expected

Where the right answers are: a column (bare or quoted) for the answer, or a named character vector c(output = "column"). Default: the columns named like the formula's outputs (an AI function) or the outcome the model was fitted on (a parsnip fit or a workflow).

metric

A function (row, prediction) returning a score (both are named lists), or a named list of such functions. Default: exact_match().

...

Settings for an AI function's calls (lm = ...), or arguments to another model's predict().

Value

An evaluation: print() it, or use generics::tidy() (the score and its 95% interval), generics::glance() and generics::augment() (every row, its answer and score).