Skip to content

News: R

functai 0.1.0 (2026-10-03)

Stages 1.2 to 5, and plugins (Python's, design/10-stages-1.2-to-5-python.md)

R now passes every contract case Python's harness reads: events/, programs/ of every kind, views/, replies/, conversations/, tools/, context/, plugins/, baked/ and saw/ of kind shown; and tools/crosslang.py checks it against Python on conversations, the disk reply cache and serving.

  • Call trees as events (contract/streaming.md): every call shows what it does (started, request, text, thinking, tool_call, tool_result, retry, done, failed, approval, approved), in one log per tree. ai_stream() watches a call while it is made (a call alone in flight is streamed from the provider), and stop_stream() cancels it. observers and journal (ai_journal(), best effort or required, with its three barriers and journal-end, settle()) keep each tree's kept form while it is written, across layers by the contract's policy (journal-policy, journal-scope, program_observers). memory_store() keeps logs by a store's rules; follower() follows them across writers; outside_view() is what a served program's caller sees.
  • Programs (ai_program()): your own code around AI functions, declared with a formula and a codebook like ai(), called as one call whose steps are the AI functions it calls; inputs bound and checked before the code runs (interface-input), outputs when it returns (interface-output); opaque() for any R value. Its version follows its code, the AI functions and programs it names, and its interface.
  • The reply cache (cache_replies): in memory, or the SQLite file every language shares ("disk"), one flight per key in a session and across processes; replicate; a progress line over a column (progress); quotes_found(); prune_calls() keeping what ratings need (every reader now reads the log folder's top-level files, and the record of a call's highest writer).
  • Conversations (ai_conversation()): turns kept in a store (memory_conversations(), folder_store(): Python's and Julia's files), branches (continue_from()), what each turn is shown (all_turns(), last_turns(), without), queueing, leases, stop_turn() from any process, request_id, earlier() in a program's turn, helpers that remember(), turn_events() from the store, ai_render() of the next turn.
  • Tools that ask first: ai_tool(.effects = "reads" | "changes"); approve (a function, "changes", "all", names or paths); a turn waits (functai_waiting), saved, and approve()/deny() from any process resume it paying for no model answer twice and running no tool twice (resume_turn() with results/rerun for a tool that may have run; abandon_turn()); a plain call refuses approval-required.
  • Plugins (ai_plugin(), on_hook(), ai_change()): seven hooks, every change recorded in the call's record; the approval plugin, compaction() and delegate() written with them; keep_entry()/entries(); load_plugin().
  • Serving (ai_serve(), ai_service(), serve_request()): the contract's routes on httpuv, keys compared in constant time, approvals to the owner or the caller; ai_remote() uses a program served by any language, one call tree across the two logs.
  • Rows that keep their context: rated() of a conversation's turns gives earlier, conversation (and helpers, sections), counts no_context; evaluate(), predict() and the optimizers ask such rows again with them, and never use one as a worked example.
  • Baking's language-neutral half: bake_examples() and export_examples() write the conversations every trainer reads (the contract's baked cases); baked() runs a student trained anywhere, served by an OpenAI-compatible server, as it was trained (baked-changed, baked-fixed, baked-derived).
  • Escalation (escalate_to, escalate_below) to another model or AI function when the first is unsure; random_search(), instruction_search(), compare(); inspect_history()/phistory(); ai_login(), ai_logins(), ai_logout().
  • Times in the log are rounded to the microsecond (they were truncated: .010000 could be written .009999).

  • A reply cut off at the token limit (contract/functions.md): it is sent again with twice max_tokens only when one was set. Without one it already had the model's whole limit (lm15's default, or the provider's own), and the old re-send with 2048 (a guessed 1024, doubled) only shrank it: it now raises at once. The refusal's hint says how many tokens went to thinking, what the limit was and whether it can be raised, and what lm15 changed in the request (a dropped thinking budget, the usual reason thinking_budget does not bound the thinking).

  • lm15's adaptations now ride on each reply of a pooled call, as lm15::complete() puts them, so the call log records them for every call.

Stage 1.1 (design/09-stage1.1-decisions.md)

  • Inputs are bound (breaking): each value is converted to its field's type when the meaning is clear (42 to text is "42", "5" to a whole number is 5), else refused before any request (interface-input, recorded). A record keeps only the members it names: a tibble with more columns than its record() sends the record's columns (before, it was refused). NA for an optional input whose type takes no null is that input left out: it takes its default (before, the row made no call). NA for a required one refuses that row, recorded (before: no call, no record). The call record holds the bound values. A default is bound when the function is defined.
  • Error messages quote the value (cut after 80 characters, ending …), except a value a log_content setting drops.
  • Defaults count in the version by what is written: defaults_to(~ Sys.Date()) (a formula) is computed at each call that leaves the input out, and counts by its code; any other default by its value (every function with a default has a new version once). A saved folder keeps a formula default's code. Signatures leave out every default.
  • The call log: a re-ask's request_hash is the hash of what it sent; outputs$calls holds every tool call of the call (it was []); returned is kept only when nothing is dropped; dropping calls keeps reasoning; process$lmcc and process$lm15.
  • Ratings: with no person named (by, or the caller's user), a rating is made under the computer's account and kept on its own, so ratings on a shared account no longer replace each other. A function defined at the top level is known by its file (the knitted document, the source()d file, the Rscript script): rated() no longer pools two notebooks' functions; rated(fn, any_file = TRUE) does.
  • The sample value of a shape whose type is a list of types is the first non-null one.

Stage 1 foundations (the contract at c1e5063)

  • The call log is format 2 (contract/calls.md): every record has program.interface, saw ([]: R's functions are shown no earlier call) and, when its content is whole, each exchange's request_hash. rated(), calls() and read_log read formats 1 and 2 and skip any other.
  • log_content per field, as a layer that only removes: the function's own, each with_ai_config() block, ai_config() and FUNCTAI_LOG_CONTENT=0 (which now wins over every setting; before, a function's own log_content beat it). c(transcript = FALSE), a named list, or c("question") (the fields that may be kept). A record not whole says what it omitted, keeps no request, reply, request hash or error message, and drops the reasoning and tool calls with any field. A name that is not one of the function's fields refuses when the function is defined (log-content-field). A one-output function's answer is result in the log; a map may call it by the formula's name too. GEPA's own calls keep no value when the function it improves drops a field.
  • defaults_to(): an input the caller may leave out; the R function's argument defaults to it, and the call sends and records it. Defaults are out of the signature and the version (functions/12).
  • ai_interface(): what an AI function, or any program in a saved folder, takes and gives, as every language describes it. ai() checks the interface when the function is defined (interface-malformed); write_ai() writes it; read_ai() checks a node's interface against its signature (saved-differs) and takes its optional inputs from it; describing a folder needs no loading (saved-no-interface for an old module node).
  • rated() pools calls by interface: calls whose data has the same shape pool, reasoning or not; a call whose inputs were not kept as data, or a right verdict on an answer not kept, is left out and counted.
  • Saved folders are checked against the contract's schema before anything else (saved-malformed); calls() gains content and saw.
  • Refusals the contract names are errors of class functai_refusal (and functai_<code>), with code and field.
  • A field name that is not an ASCII identifier (my.message) is refused when the function is defined (lmcc's signature-malformed).

After review

  • An input left out is sent with its default exactly as the interface holds it (its JSON), not through an R copy: a record's default that leaves out a member no longer gains a null for it after read_ai().
  • A default of null is sent; predict() makes one call per row of new_data even when every input is left out, and none for an empty table.
  • Values are checked on every call by the contract's vocabulary: a given input that does not fit fails its row before any request (interface-input, recorded as InterfaceError); a reply that does not fit is re-asked (parse-value), bounds, const, references, array and object rules included. A number is no longer truncated to fit a whole number, nor a value given to a choice turned into text.
  • defaults_to() casts as vctrs does (defaults_to(2.5, integer()) is refused), reads a sentence as words about the value's own type, and takes a date as text.
  • A log_content name is the field of that name: a function named like one of its inputs, or an answer named like a field FunctAI adds, no longer moves that field's rule to the answer.
  • A tool-using call's record holds outputs.calls (the value lmcc's finished turn holds: its last model step's) and its size.
  • read_saw() reads only call records of formats 1 and 2; an entry whose known keys hold values of another kind, or two different records with one id, is not guessed at.
  • A saved node without an interface is checked like any other; each probe needs its own fingerprint (saved-differs); load and describe refusals carry field; describing an optional input with no default no longer prints default null.
  • GEPA's own calls keep no value when any layer in force (a block, ai_config(), the environment) drops a field of the function it improves.
  • calls() gains omitted. Interface refusals name the first field at fault, inputs first, and the call they came from.

After the second review

  • Loading a saved function changes no value. A value is written as JSON by its field's shape alone, whatever R type the field is read as, and checked as it is: an object keeps every member it is given (a record's tibble its extra columns too), a member it lacks stays out (a required one is refused, even where it takes null), and a member a closed shape does not allow is refused. Before, a loaded record kept only its declared members and filled missing ones with null, so the same call sent another request after read_ai(), or was sent where it should have been refused.
  • write_ai() writes each field's R type in the interface's type (tibble(name = character, age = integer), list for JSON), and read_ai() reads a function R saved back with those types. Another language's function gets, per field, the R type that holds its values exactly: a record is a tibble column only when it is closed, since a tibble has no column for a member the record does not name; any other object is a list column. A record a Python dataclass wrote (open) is therefore a list column now, not a tibble.
  • In a record, NA for a member the record does not require and whose type takes no null leaves the member out (a tibble cannot); elsewhere NA is null, as before.
  • A row that makes no call because an input is missing says so in predict()'s .error; input = NULL in a direct call, for an input whose type takes no null, warns. An infinite number is refused before any request (interface-input), not by the log writer. A row refused before any request is refused once, whatever samples is.
  • A column's values are checked with each field's shape read once, not once a row: about 3.7 times faster than the first repair on a column of scalars.

After the third review

  • Another language's closed record is a tibble only when a tibble keeps whether each member is there: a member it may leave out must be one value that takes no null (NA is "left out" then, and nothing else). A record whose optional member is a list, an object or a record is a list column: before, a left-out list came back NULL and was sent back as null (refused), and a left-out record was a row of NAs, the same tibble as the record with its members null. R's own saved tibble(...) type is not believed either where it would lose that.
  • optional(record(...)) tells a null record from a record of nulls. When the record requires a member that is one value and never null, null is a row of NA (vctrs's missing row) and is sent back as null (before, a null answer given back was sent as {"name":null} and refused); when every member may be null, the field is a list column (before, a null answer and a record of nulls were the same row of NA).
  • A default is kept, shown and sent with its members in the order they were written, before and after read_ai(): ai_interface() sorted them, so a loaded function sent {"age":1,"name":"Z"} where the function it was saved from sent {"name":"Z","age":1}. A default is held as the JSON the interface holds (a number read back as JSON reads it).
  • A column named with a backslash (a\b) is written in the saved type as R writes the symbol, so R reads the tibble type back; before, the field came back a list column. Types are matched to fields by name.
  • ?ai and ?record say that every column of a tibble given for a record is sent, and to select the record's columns first when the others are not for the model.

After the fourth review

  • A record whose leaves are all null ({"child":{"x":null}}) is sent as that record, never as null, as given and as an answer given back, in a list or inside another record too. Before, any record every leaf of which was missing was sent as null where the record could be null. Now a row is read as the null record (vctrs's missing row) only when a member the record requires, one value that is never null, is NA: no record is such a row, so nothing the record admits is changed.
  • Records are closed: a record holds only the members it names. A value with another member (a tibble's id or email column given for a record()) is refused before any request (interface-input), and so is a default with one (interface-malformed) and an answer with one (the model is asked again). Before, a tibble's extra columns were sent to the model. An object shape that says additionalProperties is a shape or true, a map, and json_shape(list(type = "object")) stay open.
  • Another language's record (a Python dataclass, a TypeScript object type) is a tibble column where a tibble holds it exactly, as R's own record() is; before, it was a list column unless it said additionalProperties: false. A record that may be null is a tibble too when it requires a member that is one value and never null.
  • An optional(record(...)) held as a list column (every member may be null) prints as optional record of name (list column), not optional JSON, and takes a one-row tibble for its default, as it does for a value.
  • ?ai says the row-of-NA rule holds for any record that may be null, a json_shape() one too.

After the fifth review

  • A row of NA with a column the record does not name (tibble(name = NA, email = NA) for optional(record(name = character()))), at any depth, is refused before any request, as a row with a value in that column is, and so is a default like it; before, it was sent as null. A row of NA is null only when it names only the record's members and the member that tells null apart is NA.
  • A refusal names the member at fault for a record that may be null too: value: no member email is allowed, not … fits none of its options; a member a record does not name is named before what the members hold.
  • A json_shape() that may be null by a list of types (type = list("object", "null")) is saved; before, write_ai() failed making its sample input. Its sample value is "example text", as TypeScript and Julia make it (the contract lists no value for a list of types).

Before stage 1

The first R implementation of FunctAI, held to the same contract as the Python and TypeScript packages (../contract): every function, score, rating and saved-folder case, and a check against Python and TypeScript themselves (../tools/crosslang.py).

  • A temperature or top_p a model does not take (GPT-6, the o-series and GPT-5, Claude 5) is left out of its requests with one warning, instead of failing every row (contract/models.json, fixed_sampling).
  • ai(): a vectorised function whose body a model writes, written like a model formula: ai(team ~ message, "Which team should answer?", team = choice("shipping", "billing")). Each field is a line of a codebook: a sentence describes it, a type types it (integer(), choice(), record(), vctrs::list_of(), optional(), described()), and one left out is text. .data gives the fields their columns' types, as lm() reads them, and ~ . is every other column. Interactions, transformed columns and intercepts are refused, with the reason. Answers come back as their types, several outputs as tibble columns that mutate() splices in.
  • choice(): one of a set of answers, a factor; choice(approve = "the rules allow it", ...) tells the model what each answer means.
  • The column named in the formula is where evaluate(), the worked examples and rated() read and write the answer. In the definition a single output is still result, so the same function has the same version in every language.
  • ai_tool(lookup_order, "..."): a tool's inputs are its function's arguments, its name the function's.
  • gepa(): the instruction rewritten from the function's mistakes, as Python's GEPA does (design/04-gepa.md); ai_trials() gives the search. set_engine("functai", method = "gepa", teacher = ...) makes a tidymodels fit() learn its instruction, so resampling measures the search. Tutorials 4 and 7 use it on gpt-5.4-nano.
  • A setting given twice (.lm = "a", .lm = "b") is an error; the first used to win silently.
  • Rows run at once over curl (concurrency, default 8). A row with a missing input is NA without a call; failed rows are NA with one warning, listed by ai_problems().
  • The layouts xml (default), chat and json, templates, module = "cot", tools (ai_tool()); unreadable replies asked again in the contract's words; values checked against their types; transient provider errors re-sent.
  • evaluate() with broom's tidy(), glance(), augment(); the same scores and intervals as Python.
  • ai_model(): an AI function as a parsnip model (engine "functai"), for workflows, rsample, tune (examples = tune(), worked_examples()) and yardstick. Fitting reads the outcome's levels and picks worked examples; it calls nothing. extract_fit_engine() gives the AI function back.
  • predict() and augment() for AI functions, with tidymodels' columns: .pred_class for a choice, .pred otherwise. samples = k answers each row k times: type = "prob" gives each class's share, the class is the majority. Without it, probabilities are refused (asked directly) or NA (through parsnip): FunctAI never makes a probability up.
  • evaluate() scores any model with a predict() method (a parsnip fit, a workflow) with the same interval as an AI function.
  • vignette("getting-started"): from dplyr and ggplot2 to a first tested model, for people who have not used tidymodels.
  • vignette("tidymodels").
  • ?tickets states the shop's house rules the labels follow.
  • The call log (calls(), rate(), rated()): the same folder and records as Python and TypeScript.
  • labeled_few_shot(), bootstrap_few_shot(), with_demos(), with_instructions(), update().
  • read_ai() runs AI functions saved in Python or TypeScript; write_ai().
  • The tickets, field_notes and refunds datasets (refunds: 120 refund requests with the decision the shop's rules give, for decision models).
  • Models that measure their own probabilities (TypeSafe's Jev, lm = "jev-latest"): predict(type = "prob") and augment() give the measured .pred_<level> columns from one call a row, no votes needed; a fitted ai_model() shares those calls between its class and probability predictions; the call log's confidence is the probability the model gave its answer. Jev's calls run in parallel like any other.
  • calls() has reasoning_tokens and total_tokens: Gemini's output_tokens leave its hidden reasoning out, so a cost is total_tokens - input_tokens at the output price.