Skip to contents

Write it the way you write a model: a formula says what comes out and what goes in (team ~ message), a sentence says what the function does, and a model writes the body. The result is an ordinary, vectorised R function, team(message): call it on one value or on whole columns inside dplyr::mutate(). Each row is one model call; up to concurrency (default 8) run at once.

Usage

ai(
  .formula,
  .description = "",
  ...,
  .data = NULL,
  .name = NULL,
  .tools = NULL,
  .demos = NULL,
  .instructions = NULL,
  .defined_in = NULL
)

Arguments

.formula

outputs ~ inputs, as names joined by +.

.description

What the function does, in a sentence or a paragraph: the model's instructions.

...

The fields, as above; and settings, as dotted names (.lm = "claude-haiku-4-5", .temperature = 0, .adapter = "json", .module = "cot", ...: see ai_config()).

.data

A data frame whose columns give the fields their types (a factor becomes a choice of its levels). Only its column types are read, never its rows.

.name

The function's name. The model reads it ("Function: team"), and the call log files calls under it. Default: the output's name, when there is one.

.tools

Functions the model may call, from ai_tool().

.demos

Worked examples: a data frame with input and output columns.

.instructions

An improved instruction that replaces the written one.

.defined_in

The code module the call log files calls under (default "__main__", as a Python notebook's).

Value

An AI function (class functai_fn).

Details

The formula. Outputs on the left, inputs on the right, each a name, joined by +: team ~ message, summary + urgent ~ message, decision ~ message + price + final_sale. With .data, . on the right is every other column, as in stats::lm(). One output: the function returns a vector and is named after it (team). Several: it returns a tibble, one column each, which dplyr::mutate() splices in; name it with .name.

The fields (..., named like the formula's), are a codebook:

What it takes and gives is its interface (ai_interface()), checked when the function is defined, as every language checks one: field names are ASCII identifiers (order_id, not order.id), and a default fits its type. A function that breaks these rules is refused then (interface-malformed), not later by whatever reads it.

Values are bound on every call (contract/programs.md, "Binding a call's inputs"), as every language binds them: a value is converted to its field's type when the meaning is clear, and refused otherwise. Text takes a number (42 is "42"), TRUE/FALSE ("true", "false"), a date (its text), a list or a one-row tibble (JSON indented by two spaces); a whole number takes 5 or "5", not 2.5; a number takes "2.5"; TRUE/FALSE take nothing else. A record keeps only the members it names: a tibble with more columns than the record() names (an id, an e-mail address) sends the record's columns, and the others are dropped. A value that does not bind fails that row before any request (interface-input, a functai_interface_input condition naming the field, its message quoting the value unless the log drops that field), and its call record says so; a reply whose value does not fit is unreadable, and the model is asked again. The call record holds the bound values. An input left out is sent with its default exactly as the interface holds it; a default written as a formula (defaults_to(~ Sys.Date())) is computed at each call that leaves it out.

Missing values. NA (or NULL) is JSON's null. An input whose type takes null (an optional() type, a json_shape() that allows null) is sent null. An optional input whose type takes no null is left out: it is sent with its default, as a table's gap takes the default. A required one fails that row before any request (interface-input): its answer is NA, and predict()'s .error says why. Inside a record (a row of a tibble column, or a named list), NA is null too, except for a member the record does not require and whose type takes no null: a tibble cannot leave a member out, so NA there leaves it out, as an answer that leaves it out comes back NA.

What the call log keeps is .log_content (see ai_config()): TRUE, FALSE, c(transcript = FALSE), or the names of the fields it may keep, c("question"). It only removes: a block around the call, or ai_config(), may drop more; nothing brings back what one of them drops. A name that is not one of the function's fields is refused (log-content-field). A one-output function's answer is result in the log, as in every language; log_content may call it by the formula's name too (c(team = FALSE)).

Examples

mood <- ai(mood ~ review, "How does the customer feel about what they bought?",
  mood = choice("happy", "unhappy", "mixed"))
mood
#> <ai function> mood ~ review
#>   How does the customer feel about what they bought?
#>   review  text
#>   mood    one of happy, unhappy, mixed
#>   model: (the default)

triage <- ai(summary + urgent ~ message, "Read the support ticket.",
  message = "the customer's own words",
  summary = "one sentence, no names",
  urgent = logical(),
  .name = "triage")

# types from a table's columns: `decision` is a factor, `price` a number
refund <- ai(decision ~ message + price + final_sale, "Should the shop refund this request?",
  .data = refunds)
if (FALSE) { # \dontrun{
mood("It broke after one day.")
tickets |> dplyr::mutate(triage(message))
} # }

# an input the caller may leave out
reply <- ai(reply ~ message + tone, "Answer the customer.",
  tone = defaults_to("kind", choice("kind", "brief")))
args(reply)
#> function (message, tone = "kind") 
#> NULL