Skip to content

Programs of several AI functions: a multi-hop fact checker

A multi-hop fact checker as one @module: evaluated, optimized, run on a table.

A @module is a plain Python function that calls AI functions: loops, ifs and ordinary helpers included. It is called, evaluated, optimized and saved as one program. This one checks a claim by searching a small library in hops, taking notes, then deciding.

Every output below is a real reply. This page is a notebook: open it in Chattering and run it, or run it all with python/.venv/bin/python tools/docs.py run python/examples/modules/README.md.

import functai
functai.configure(lm="gpt-4.1-mini", temperature=0)

from functai import ai, module

A search engine (plain Python)

LIBRARY = [
    "The Eiffel Tower was completed in 1889 for the World's Fair in Paris.",
    "Gustave Eiffel's company designed and built the Eiffel Tower.",
    "The Statue of Liberty's internal frame was designed by Gustave Eiffel.",
    "The Statue of Liberty was a gift from France to the United States in 1886.",
    "Mount Everest, at 8,849 m, is the highest mountain above sea level.",
    "K2 is the second-highest mountain on Earth, at 8,611 m.",
    "The Great Wall of China is not visible to the naked eye from orbit.",
    "Marie Curie won Nobel Prizes in both physics (1903) and chemistry (1911).",
]

def search(query: str, k: int = 2) -> list[str]:
    """The k passages sharing the most words with the query."""
    words = set(query.lower().split())
    return sorted(LIBRARY, key=lambda p: -len(words & set(p.lower().split())))[:k]

Three AI functions and the program

from typing import Literal

Verdict = Literal["supported", "refuted", "not enough info"]

@ai
def next_query(claim: str, notes: list[str]) -> str:
    """A short search query for the fact still missing to check the claim."""
    ...

@ai
def take_notes(claim: str, notes: list[str], passages: list[str]) -> list[str]:
    """The notes, plus what the passages say that bears on the claim."""
    ...

@ai
def decide(claim: str, notes: list[str]) -> Verdict:
    """Is the claim supported or refuted by the notes?"""
    ...

@module
def check_claim(claim: str, hops: int = 2) -> Verdict:
    notes: list[str] = []
    for _ in range(hops):
        notes = take_notes(claim, notes, search(next_query(claim, notes)))
    return decide(claim, notes)

check_claim("The man who designed the Eiffel Tower also worked on the Statue of Liberty.")
'supported'

Evaluating the whole program

Rows name the module’s inputs (claim) and the expected result. The metric sees what the module returned; the table’s tokens add up every model call the program made.

dev = [
    {"claim": "The Eiffel Tower was finished before the Statue of Liberty was given to the US.", "result": "refuted"},
    {"claim": "K2 is taller than Mount Everest.", "result": "refuted"},
    {"claim": "Marie Curie won two Nobel Prizes in different sciences.", "result": "supported"},
    {"claim": "The Great Wall of China can be seen from orbit with the naked eye.", "result": "refuted"},
    {"claim": "Gustave Eiffel's company built a tower completed in 1889.", "result": "supported"},
    {"claim": "The Eiffel Tower was built for a World's Fair.", "result": "supported"},
]

ev = functai.evaluate(check_claim, dev, num_threads=6)
ev
Evaluation(check_claim, 6 examples: exact_match 0.83 [0.44, 0.97])
from dpyr import col

ev.table.select(col.claim, col.result, col.pred_result, col.input_tokens, col.seconds)
# dpyr dataframe · source: polars · showing 6 of 6 rows
┌─────────────────────────────────────────────────────────────────────────────────┬───────────┬─────────────────┬──────────────┬──────────┐
│ claim                                                                           ┆ result    ┆ pred_result     ┆ input_tokens ┆ seconds  │
│ ---                                                                             ┆ ---       ┆ ---             ┆ ---          ┆ ---      │
│ str                                                                             ┆ str       ┆ str             ┆ i64          ┆ f64      │
╞═════════════════════════════════════════════════════════════════════════════════╪═══════════╪═════════════════╪══════════════╪══════════╡
│ The Eiffel Tower was finished before the Statue of Liberty was given to the US. ┆ refuted   ┆ refuted         ┆ 660          ┆ 4.32156  │
│ K2 is taller than Mount Everest.                                                ┆ refuted   ┆ not enough info ┆ 461          ┆ 4.081073 │
│ Marie Curie won two Nobel Prizes in different sciences.                         ┆ supported ┆ supported       ┆ 624          ┆ 4.408497 │
│ The Great Wall of China can be seen from orbit with the naked eye.              ┆ refuted   ┆ refuted         ┆ 595          ┆ 4.334084 │
│ Gustave Eiffel's company built a tower completed in 1889.                       ┆ supported ┆ supported       ┆ 590          ┆ 4.308764 │
│ The Eiffel Tower was built for a World's Fair.                                  ┆ supported ┆ supported       ┆ 586          ┆ 4.459435 │
└─────────────────────────────────────────────────────────────────────────────────┴───────────┴─────────────────┴──────────────┴──────────┘

Optimizing it

Optimizing a module tunes every AI function it calls, against the one metric on the program’s output. A run the metric accepts becomes a worked example (a “demo”) for each AI function it went through.

train = [
    {"claim": "Mount Everest is higher than 8,000 m.", "result": "supported"},
    {"claim": "The Statue of Liberty was a gift from Germany.", "result": "refuted"},
    {"claim": "Marie Curie's chemistry Nobel came before her physics one.", "result": "refuted"},
    {"claim": "Eiffel's company also designed part of the Statue of Liberty.", "result": "supported"},
]

taught = check_claim.opt(train)                   # an improved copy
{name: len(state.demos) for name, state in taught.state().items()}
{'take_notes': 4, 'next_query': 4, 'decide': 2}
after = functai.evaluate(taught, dev, num_threads=6)
functai.compare(ev, after)
# dpyr dataframe · source: polars · showing 1 of 1 rows
┌─────────────┬──────────┬──────────┬──────┬─────┬──────┬────────┬───────┬──────┬─────┐
│ metric      ┆ before   ┆ after    ┆ diff ┆ low ┆ high ┆ better ┆ worse ┆ same ┆ n   │
│ ---         ┆ ---      ┆ ---      ┆ ---  ┆ --- ┆ ---  ┆ ---    ┆ ---   ┆ ---  ┆ --- │
│ str         ┆ f64      ┆ f64      ┆ f64  ┆ f64 ┆ f64  ┆ i64    ┆ i64   ┆ i64  ┆ i64 │
╞═════════════╪══════════╪══════════╪══════╪═════╪══════╪════════╪═══════╪══════╪═════╡
│ exact_match ┆ 0.833333 ┆ 0.833333 ┆ 0.0  ┆ 0.0 ┆ 0.0  ┆ 0      ┆ 0     ┆ 6    ┆ 6   │
└─────────────┴──────────┴──────────┴──────┴─────┴──────┴────────┴───────┴──────┴─────┘

check_claim itself, and the AI functions it calls, are unchanged; taught.save("check_claim.json") keeps the copy's instructions and demos (check_claim.load(...) is a copy running with them).

On a table

A module with a return type works on columns like an AI function:

from dpyr import read

read(dev).mutate(verdict=check_claim(col.claim)).select(col.claim, col.verdict)
# dpyr dataframe · source: polars · showing 6 of 6 rows
┌─────────────────────────────────────────────────────────────────────────────────┬─────────────────┐
│ claim                                                                           ┆ verdict         │
│ ---                                                                             ┆ ---             │
│ str                                                                             ┆ str             │
╞═════════════════════════════════════════════════════════════════════════════════╪═════════════════╡
│ The Eiffel Tower was finished before the Statue of Liberty was given to the US. ┆ refuted         │
│ K2 is taller than Mount Everest.                                                ┆ not enough info │
│ Marie Curie won two Nobel Prizes in different sciences.                         ┆ supported       │
│ The Great Wall of China can be seen from orbit with the naked eye.              ┆ refuted         │
│ Gustave Eiffel's company built a tower completed in 1889.                       ┆ supported       │
│ The Eiffel Tower was built for a World's Fair.                                  ┆ supported       │
└─────────────────────────────────────────────────────────────────────────────────┴─────────────────┘