Skip to content

Evaluation

Evaluation(target, rows, runs, metrics, values, run_id)

The result of evaluate: a score, its uncertainty, and every answer.

ev.scores(metric) gives each row's value for a metric (the first by default); ev.write("run.parquet") saves the table.

Attributes

Name Type Description
score float The first metric's mean, from 0 to 1. A failed row counts 0.
summary dpyr dataframe One row per metric: mean, the 95% interval low to high, n, and how many rows failed.
table dpyr dataframe One row per example: the data, pred_<output> for each output, each metric, error, seconds, input_tokens, output_tokens, reasoning_tokens, total_tokens, model and run.
predictions list Each row's Prediction (None where it failed), with its turns, tokens and repairs.
errors list (row number, message) for each row that failed.

See Also

Methods

Name Description
scores One metric's value per row, a failed row counting 0.
write Save the table (.parquet keeps every type; also .jsonl, .arrow, ...).

scores

Evaluation.scores(metric=None)

One metric's value per row, a failed row counting 0.

write

Evaluation.write(path)

Save the table (.parquet keeps every type; also .jsonl, .arrow, ...).