Watch it being written¶
Show the answer as the model writes it: text, reasoning, tool calls, records filling in, in a notebook, a script or a web app.
import functai
functai.configure(lm="gpt-4.1-mini", temperature=0) # the model behind every output on this page
from functai import ai, _ai
A long answer takes seconds to write. Calling the function waits for all
of it; .stream(...) makes the same call and lets you watch it being
written.
Print it as it arrives¶
@ai
def story(topic: str) -> str:
"""A four-sentence story about the topic."""
...
story.stream("a lighthouse keeper's cat").show()
Every evening, the lighthouse keeper's cat would curl up by the warm lantern, watching the waves crash below. One stormy night, the cat's sharp eyes noticed a ship struggling against the fierce winds. It meowed loudly, alerting the keeper to adjust the light just in time to guide the vessel safely to shore. From that day on, the cat was known as the lighthouse's silent guardian, watching over both sea and land.
.show() prints the answer as it arrives and waits for the end. In a
notebook, the words appear one after another. To do something else with
each piece, iterate:
pieces = []
for piece in story.stream("a kettle that sings"):
pieces.append(piece) # or print(piece, end="", flush=True)
len(pieces), "".join(pieces)[:60]
(81, 'Every morning, the old kettle on the stove would sing a chee')
The same call¶
A stream is not a different kind of call. It asks the model for the same thing, reads the answer the same way, retries and uses tools the same way, and ends with the same value:
s = story.stream("a kettle that sings")
s.result
'Every morning, the old kettle on the stove would sing a cheerful tune as it boiled water. Its melodic whistle brought joy to the entire kitchen, waking everyone with a smile. One day, the family discovered that the kettle’s song changed with the weather, humming softly on rainy days and loudly on sunny ones. From then on, the singing kettle became their magical weather forecast and a beloved member of the household.'
s.result waits for the end and gives what story(...) returns (or
raises what it raises). s.prediction is what predict returns, with
the tokens and every message. With the call log on, the
call gets the same line, which also records how long the first word
took.
The call starts as soon as .stream(...) returns, in the background: you
can start several and read them as they come.
Reasoning, then the answer¶
A function that writes a reasoning before its answer shows both.
.show() labels them:
@ai
def solve(problem: str) -> float:
"""Solve the word problem."""
reasoning: str = _ai["Step by step, briefly."]
return _ai
solve.stream("3 pencils cost $1.20. How much do 10 cost?").show()
reasoning: First, find the cost of one pencil by dividing the total cost by the number of pencils: $1.20 ÷ 3 = $0.40 per pencil. Then, multiply the cost per pencil by 10 to find the cost of 10 pencils: $0.40 × 10 = $4.00.
result: 4.00
Iterating the stream gives the answer's text only. s.events() gives
everything, in order, as small objects with a kind:
s = solve.stream("A train leaves at 3:40 and arrives at 5:15. How long is the trip, in minutes?")
pieces = {}
for event in s.events():
if event.kind == "text":
pieces.setdefault(event.field, []).append(event.text)
else:
print(event.kind)
{field: (len(texts), texts[:6]) for field, texts in pieces.items()} # how many pieces, the first ones
started
request
done
{'reasoning': (49, ['The', ' train', ' leaves', ' at', ' 3', ':']), 'result': (1, ['95'])}
kind |
what happened |
|---|---|
started |
a call began (inputs) |
text |
a piece of an output (field, text; answer when it is the answer) |
thinking |
a piece of a thinking model's own thinking |
tool_call, tool_result |
the model used a tool, and what the tool said |
retry |
the model is asked again (reason) |
done, failed |
a call ended (value, or error) |
Tools¶
Tool calls are shown when the model has finished asking, and their results when the tool ran:
def get_weather(city: str) -> str:
"""Current weather for a city."""
return {"Oslo": "Light snow, -2C.", "Lima": "Cloudy, 18C."}.get(city, "Unknown.")
@ai(tools=[get_weather])
def assistant(question: str) -> str:
"""Answer; use a tool when you need facts."""
...
assistant.stream("Should I pack a coat for Oslo or for Lima?").show()
→ get_weather(city='Oslo')
← Light snow, -2C.
→ get_weather(city='Lima')
← Cloudy, 18C.
Oslo currently has light snow and a temperature of -2°C, so you should definitely pack a coat for Oslo. Lima, on the other hand, is cloudy with a mild temperature of 18°C, so a coat is not necessary there.
Records and lists, filling in¶
When the answer is a record or a list, s.partial is the answer so far,
read from what has been written: the complete values, and each text as
far as it goes. A number appears only once it is complete, so what you
see is never a wrong number, only an unfinished list:
from dataclasses import dataclass
@dataclass
class Person:
name: str
city: str | None # None when the text does not say
@ai
def people(text: str) -> list[Person]:
"""Everyone the text mentions."""
...
s = people.stream("Ada wrote from London to Grace in Arlington; Alan answered from Manchester, and Kurt said nothing.")
seen = []
for piece in s:
if s.partial and s.partial not in seen:
seen.append(s.partial)
for view in seen[::3]:
print(view)
s.result
[{}]
[{'name': 'Ada', 'city': 'London'}, {}]
[{'name': 'Ada', 'city': 'London'}, {'name': 'Grace', 'city': 'Arlington'}]
[{'name': 'Ada', 'city': 'London'}, {'name': 'Grace', 'city': 'Arlington'}, {'name': 'Alan'}]
[{'name': 'Ada', 'city': 'London'}, {'name': 'Grace', 'city': 'Arlington'}, {'name': 'Alan', 'city': 'Manchester'}, {}]
[{'name': 'Ada', 'city': 'London'}, {'name': 'Grace', 'city': 'Arlington'}, {'name': 'Alan', 'city': 'Manchester'}, {'name': 'Kurt', 'city': None}]
[Person(name='Ada', city='London'), Person(name='Grace', city='Arlington'), Person(name='Alan', city='Manchester'), Person(name='Kurt', city=None)]
s.partial is plain data, provisional; s.result is the checked,
typed answer. s.text is the answer's text so far, and s.fields every
output's.
When the model is asked again¶
If a reply cannot be read (a choice outside the allowed ones, a reply cut
off), FunctAI asks the model again, as a normal call does. The stream
shows a retry event, and the answer starts again: s.text and
s.partial restart with it, and .show() says so. Pieces already
handed to a for loop cannot be taken back, so "".join(pieces) is the
text as shown; the answer is s.result.
Modules¶
A module's stream shows every AI function it calls, as it calls them:
from functai import module
@ai
def draft(topic: str) -> str:
"""A paragraph about the topic."""
...
@ai
def shorten(text: str) -> str:
"""The text in at most twelve words."""
...
@module
def blurb(topic: str) -> str:
return shorten(draft(topic))
blurb.stream("tide pools").show()
▸ draft
Tide pools are fascinating coastal ecosystems found in the rocky intertidal zones where seawater collects during low tide. These pools serve as temporary habitats for a diverse array of marine life, including sea stars, anemones, crabs, and small fish. The unique conditions of tide pools, such as fluctuating water levels, temperature, and salinity, create a challenging environment that supports specially adapted organisms. Exploring tide pools offers valuable insights into marine biodiversity and the delicate balance of coastal ecosystems.
▸ shorten
Tide pools are diverse, temporary coastal habitats with unique marine life.
s.text_of(shorten) gives one function's answer as it is written, and
s.result what the module returned.
Stopping¶
Closing a stream stops the call: the model stops at its next piece, and
the call ends with functai.Cancelled (in the call log too). Leaving a
with block closes it:
with story.stream("an endless staircase") as s:
for i, piece in enumerate(s):
if i == 10:
break # the with block closes it
s
<Stream story: closing>
Ctrl-C while watching stops it too. A stream you never close runs to its end, like a call. The provider may still bill the words it wrote before it stopped.
In async code, and in a web app¶
Streams work with async for, and await s gives the result without
blocking the event loop:
import asyncio
async def main():
s = story.stream("a kettle that sings")
n = 0
async for piece in s:
n += 1
return n, (await s)[:40]
asyncio.run(main()) # in Jupyter, where an event loop is already running: await main()
(82, 'Every morning, the old kettle on the sto')
Each event has .to_dict(), plain JSON data, to send to a browser (as
server-sent events, or over a websocket). If the consumer goes away (the
browser tab closes and the server cancels its task), the call stops. The
events are a written contract
(contract/streaming.md),
the same for FunctAI in other languages.
What a caller may see¶
The events above are the full view: everything, values included. Two narrower views are made from them, event by event:
- kept: what may be written down, as the
log_contentsettings allow (a field kept out of the log is kept out of these events too); - outside: what someone who only sees the program's boundary may see, a customer of a served program, say: the program's answer and its text as it is written, approvals addressed to them, and the end. Never a helper's answer, a tool call or its result, a model's thinking, why a request was retried, or an error's message.
A module's answer is usually one of its helpers' answers. answer_from
says which, so the outside view shows that text being written, as the
module's own:
from typing import Literal
@ai
def topic(message: str) -> Literal["billing", "shipping", "other"]:
"""What the message is about."""
...
@ai
def answer(message: str, topic: str) -> str:
"""Answer the customer in one short sentence."""
...
@module(answer_from=answer)
def support(message: str) -> str:
return answer(message, topic(message))
s = support.stream("Where is my parcel B-2210?")
s.result
full = [(e.kind, e.function) for e in s.events() if e.kind != "text"]
outside = [(e["kind"], e["function"]) for e in s.events(view="outside") if e["kind"] != "text"]
full, outside
([('started', 'support'), ('started', 'topic'), ('request', 'topic'), ('done', 'topic'), ('started', 'answer'), ('request', 'answer'), ('done', 'answer'), ('done', 'support')], [('started', 'support'), ('request', 'support'), ('request', 'support'), ('done', 'support')])
A view's events are the contract's JSON, ready to send to a browser.
Every event, kept as it happens¶
A stream is one call, watched by whoever made it. To see every call a process makes (a dashboard, an audit trail), give observers: each gets every event of every call tree, in the kept view, as plain dicts. A list collects them; a function is called with each, in a thread of its own, so a slow observer never slows a call:
seen = []
with functai.configure(observers=[seen]):
support("I was charged twice for order B-2210.")
len(seen), sorted({e["kind"] for e in seen})
(31, ['done', 'request', 'started', 'text'])
functai.flush() waits until every function observer has caught up,
and every best-effort journal has written what it was given (at exit,
FunctAI waits for them at most two seconds).
A journal keeps whole call trees in a store while they run, so
another process can follow a call, or find what a crashed one did.
functai.MemoryStore is a store in this process's memory; any object
that keeps events by the contract's rules is one too (functai.Store
says what it must do):
store = functai.MemoryStore()
with functai.configure(journal=store):
support("The kettle lid doesn't close.")
functai.flush() # a best-effort journal writes in the background
tree = store.trees()[-1]
[e["kind"] for e in store.read(tree, None) if e["kind"] != "text"]
['started', 'started', 'request', 'done', 'started', 'request', 'done', 'done']
A journal is best effort by default: if the store fails, FunctAI
warns once and the call goes on. functai.Journal(store, required=True)
makes the call wait until the store has confirmed what happened so far,
at three moments: when it starts, before each tool runs, and before it
returns. If the store does not confirm, the call stops before the code or
the tool runs (JournalError, journal-barrier), or the caller is told
the end was not confirmed (journal-end, which still holds the call's
outcome). A program cannot replace or remove the journal its host set
(journal-policy).
A reader that follows a log, live or later, is a functai.Follower: it
takes events in any order and from any source (duplicates, a writer that
took over after a crash), and keeps each tree's state:
follower = functai.Follower()
for event in store.read(tree, None):
follower.receive(event)
follower.state(tree)["finished"], [c["fields"] for c in follower.state(tree)["calls"].values()]
(True, [{}, {'result': 'other'}, {'result': 'Please check if there is any obstruction or misalignment preventing the lid from closing properly.'}])
What streams, and what doesn't¶
- Every provider that lm15 can stream from streams. A model or client that can't (a baked model, a judgment-only provider) answers whole, and the stream shows the answer in one piece.
- A reply from the reply cache is shown in one piece.
- With
adapter="json", each output appears whole at the end of the reply: the JSON reader does not yet read a reply in pieces.