Runs a model on every row of data and scores each answer against the
right one, with a 95% interval. The model is an AI function, a fitted
parsnip model, a fitted workflow, or anything with a predict() method
returning .pred_class or .pred: the same score, the same interval, so
an AI function and a classical model compare on equal terms. A row whose
call failed scores 0. An AI function's calls are logged with
caller.evaluation set, so they are not mistaken for use.
Arguments
- fn
A model: an AI function, a parsnip fit, a workflow, ...
- data
A data frame: the model's inputs, and the right answers.
- expected
Where the right answers are: a column (bare or quoted) for the answer, or a named character vector
c(output = "column"). Default: the columns named like the formula's outputs (an AI function) or the outcome the model was fitted on (a parsnip fit or a workflow).- metric
A function
(row, prediction)returning a score (both are named lists), or a named list of such functions. Default:exact_match().- ...
Settings for an AI function's calls (
lm = ...), or arguments to another model'spredict().
Value
An evaluation: print() it, or use generics::tidy() (the score
and its 95% interval), generics::glance() and generics::augment()
(every row, its answer and score).