Skip to content

Multi-step programs

@module: ordinary Python that calls AI functions, measured, optimized and saved as one program.

import functai
functai.configure(lm="gpt-4.1-mini", temperature=0)   # the model behind every output on this page
from functai import ai, _ai

Real tasks rarely fit in one model call. You search, then take notes, then decide; you draft, then check. In functai the program that ties those calls together is plain Python: loops, ifs, helpers. Put @module on it and it becomes one program you can evaluate, optimize and save.

A research loop

Two AI functions and an ordinary search function:

@ai
def next_query(claim: str, notes: list[str]) -> str:
    """A search query that would help check the claim, given the notes so far."""
    ...

@ai
def take_notes(claim: str, notes: list[str], documents: list[str]) -> list[str]:
    """The notes, extended with what the documents say about the claim."""
    ...

LIBRARY = {
    "eiffel": "The Eiffel Tower is a wrought-iron tower in Paris, completed in 1889.",
    "k2": "K2 is the second-highest mountain on Earth, on the China–Pakistan border.",
    "everest": "Mount Everest lies on the border between Nepal and China.",
}

def search(query: str) -> list[str]:
    """Your retriever: a vector store, a search API, ..."""
    words = query.lower().split()
    return [text for key, text in LIBRARY.items() if any(key in w for w in words)] or ["No results."]

The program calls them in a loop, and a third AI function decides:

from typing import Literal
from functai import module

@ai
def verdict(claim: str, notes: list[str]) -> Literal["true", "false", "unknown"]:
    """Is the claim true, according to the notes?"""
    ...

@module
def fact_check(claim: str, hops: int = 2) -> Literal["true", "false", "unknown"]:
    notes: list[str] = []
    for _ in range(hops):
        query = next_query(claim, notes)
        notes = take_notes(claim, notes, search(query))
    return verdict(claim, notes)

fact_check("K2 is in Nepal.")
'false'

phistory(5) shows the calls it made, in order:

print(functai.phistory(5)[:1500], "…")
[2026-10-02T09:58:49] next_query → gpt-4.1-mini

System message:

Function: next_query

A search query that would help check the claim, given the notes so far.

Reply in exactly this form:
<result>
...
</result>


User message:

<claim>
K2 is in Nepal.
</claim>
<notes>
[]
</notes>


Response:

<result>
Is K2 located in Nepal?
</result>

(finish: stop; tokens in 64, out 14)

────────────────────────────────────────────────────────────

[2026-10-02T09:58:50] take_notes → gpt-4.1-mini

System message:

Function: take_notes

The notes, extended with what the documents say about the claim.

Reply in exactly this form:
<result>
JSON matching this schema: {"type": "array", "items": {"type": "string"}}
</result>


User message:

<claim>
K2 is in Nepal.
</claim>
<notes>
[]
</notes>
<documents>
[
  "K2 is the second-highest mountain on Earth, on the China–Pakistan border."
]
</documents>


Response:

<result>
["K2 is not in Nepal; it is located on the China–Pakistan border."]
</result>

(finish: stop; tokens in 109, out 25)

────────────────────────────────────────────────────────────

[2026-10-02T09:58:51] next_query → gpt-4.1-mini

System message:

Function: next_query

A search query that would help check the claim, given the notes so far.

Reply in exactly this form:
<result>
...
</result>


User message:

<claim>
K2 is in Nepal.
</claim>
<notes>
[
  "K2 is not in Nepal; it is located on the China–Pakistan border."
]
</notes>


Response:

<result>
K2 location China Pakistan border
 …

Measured and improved as one

A module is evaluated like a single function. The metric sees what the module returned (pred_result in the table), and the tokens of every inner call are added up per row.

claims = [
    {"claim": "The Eiffel Tower is in Paris.", "result": "true"},
    {"claim": "K2 is in Nepal.", "result": "false"},
    {"claim": "Mount Everest is on the border of Nepal.", "result": "true"},
    {"claim": "The Eiffel Tower was finished in 1920.", "result": "false"},
]

ev = functai.evaluate(fact_check, claims, num_threads=4)
ev
Evaluation(fact_check, 4 examples: exact_match 1.00 [0.51, 1.00])

Optimizing a module improves each AI function inside it: every run the metric accepts gives a worked example to each function it went through.

better = fact_check.opt(claims, call_defaults=dict(hops=1))   # an improved copy

Why a module and not a plain function?

A plain Python function that calls AI functions works; you just can't treat it as one thing. @module adds:

  • evaluate(program, data) and program.opt(rows) (an improved copy) for the whole program;
  • a typed signature (claim: str → Literal[...]), so it can be saved, run on a table, and checked;
  • functai.check(program) follows every function, tool and constant it reaches, so saving takes all of it.