Skip to content

Building a knowledge graph, one chunk at a time

A knowledge graph built chunk by chunk with pydantic models, then queried.

Text arrives in chunks; we want one knowledge graph out of all of them. An AI function reads a chunk together with the graph so far and returns only what is new; plain Python merges it in. Pydantic models are the contract on both sides.

Every output below is a real reply. This page is a notebook: open it in Chattering and run it, or run it all with python/.venv/bin/python tools/docs.py run python/examples/graph_rag/README.md.

import functai
functai.configure(lm="gpt-4.1-mini", temperature=0)

from functai import ai

The graph

from pydantic import BaseModel

class Node(BaseModel, frozen=True):
    id: int
    label: str

class Edge(BaseModel, frozen=True):
    source: int   # a node id
    target: int   # a node id
    label: str    # the relation, as a short verb phrase

class KnowledgeGraph(BaseModel):
    nodes: list[Node] = []
    edges: list[Edge] = []

    def merge(self, other: "KnowledgeGraph") -> "KnowledgeGraph":
        """This graph plus the other's new nodes and edges."""
        nodes = list(dict.fromkeys(self.nodes + other.nodes))
        edges = list(dict.fromkeys(self.edges + other.edges))
        return KnowledgeGraph(nodes=nodes, edges=edges)

The AI function

The graph so far is an input like any other: it is shown to the model as JSON, and the reply is read back as a KnowledgeGraph.

@ai
def new_facts(text: str, graph: KnowledgeGraph) -> KnowledgeGraph:
    """Extract the entities and relations in the text that are not in the
    graph yet. Reuse the ids of nodes the graph already has; give new nodes
    ids that are not taken."""
    ...

Building it

chunks = [
    "Jason knows a lot about quantum mechanics. He is a physicist and a professor.",
    "Professors teach at universities.",
    "Sarah knows Jason and is a student of his.",
    "Sarah studies at the University of Toronto, which is in Canada.",
]

graph = KnowledgeGraph()
for chunk in chunks:
    graph = graph.merge(new_facts(chunk, graph))

len(graph.nodes), len(graph.edges)
(8, 8)

Looking at it

A Mermaid diagram renders directly on GitHub:

def mermaid(g: KnowledgeGraph) -> str:
    lines = ["graph LR"]
    lines += [f'  n{n.id}["{n.label}"]' for n in g.nodes]
    lines += [f'  n{e.source} -- "{e.label}" --> n{e.target}' for e in g.edges]
    return "\n".join(lines)
print("```mermaid\n" + mermaid(graph) + "\n```")
graph LR
  n1["Jason"]
  n2["quantum mechanics"]
  n3["physicist"]
  n4["professor"]
  n5["university"]
  n6["Sarah"]
  n7["University of Toronto"]
  n8["Canada"]
  n1 -- "knows about" --> n2
  n1 -- "is a" --> n3
  n1 -- "is a" --> n4
  n4 -- "teach at" --> n5
  n6 -- "knows" --> n1
  n6 -- "student of" --> n1
  n6 -- "studies at" --> n7
  n7 -- "is in" --> n8

Asking it questions

The graph is data now. Another AI function can answer from it, with no access to the original text:

@ai
def ask(graph: KnowledgeGraph, question: str) -> str:
    """Answer from the graph only. Say so when the graph doesn't tell."""
    ...

ask(graph, "In which country does Jason's student study?")
"Jason's student, Sarah, studies at the University of Toronto, which is in Canada."
ask(graph, "How old is Sarah?")
"The graph doesn't tell how old Sarah is."