Skip to content

bake.Run

bake.Run(folder)

A training run: a folder, a process of its own, and a model at the end.

What bake(..., wait=False) returns at once (and Plan.run, functai.bake.runs(), functai.bake.run(folder)). The run is its own process, so it outlives the notebook or terminal that started it; running the same bake again with the same plan finds its folder and resumes it (or finds it done).

run = summarize.bake(rows, wait=False)   # returns at once
run.metrics()                            # the loss curve so far, as a table
run.stop(); run.resume()                 # from the last checkpoint
baked = run.wait()                       # the model, when it is done
run.checkpoint(1200)                     # any checkpoint as a model

Its folder holds plan.json (every resolved setting: what makes resuming exact), examples.parquet (the training conversations), run.json (its state, where it trains, its process, its attempts), metrics.jsonl (one line per logged step: step, tokens, loss, learning rate, validation loss, seconds), log.txt (the trainer's own output), checkpoints/ and, when done, baked/ (the model).

Attributes

Name Description
plan Every setting the run was decided with (its plan.json).
state planned, starting, running, stopping, stopped (by stop(), or its
status run.json, with a process that ended without saying so shown as such.
where Where it trains: here, tinker, prime or export.

Methods

Name Description
baked The model this run made (when done).
checkpoint A checkpoint as a usable model (the last one by default). Models from
checkpoints Steps with a checkpoint, oldest first.
log The last lines lines of the trainer's own output (log.txt): read it when a run fails.
log_metric Append one line to metrics.jsonl (what trainers call).
metrics The logged steps as a dpyr table (a list of dicts without dpyr).
progress Where it is now: state, step of steps, the last loss and eval_loss,
records The lines of metrics.jsonl so far, as dicts (metrics() gives them as a table).
resume Continue a stopped or crashed run from its last checkpoint.
start Start (or resume) the run in its own process; process=False runs it
stop Ask the run to stop at its next step (it saves a checkpoint first).
stop_requested Whether stop() was asked (what the trainer checks at each step).
update Change run.json (the supervisor's; also stop).
wait Wait for the end; show progress; return the model.

baked

bake.Run.baked()

The model this run made (when done).

checkpoint

bake.Run.checkpoint(step=None)

A checkpoint as a usable model (the last one by default). Models from the constant phase of the schedule are not decayed: compare them with each other, not with the final model.

checkpoints

bake.Run.checkpoints()

Steps with a checkpoint, oldest first.

log

bake.Run.log(lines=40)

The last lines lines of the trainer's own output (log.txt): read it when a run fails.

log_metric

bake.Run.log_metric(**fields)

Append one line to metrics.jsonl (what trainers call).

metrics

bake.Run.metrics()

The logged steps as a dpyr table (a list of dicts without dpyr).

progress

bake.Run.progress()

Where it is now: state, step of steps, the last loss and eval_loss, seconds so far and eta (seconds left, estimated).

records

bake.Run.records()

The lines of metrics.jsonl so far, as dicts (metrics() gives them as a table).

resume

bake.Run.resume()

Continue a stopped or crashed run from its last checkpoint.

start

bake.Run.start(process=True)

Start (or resume) the run in its own process; process=False runs it in a thread of this one (ends with it).

stop

bake.Run.stop(wait=True, timeout=600)

Ask the run to stop at its next step (it saves a checkpoint first).

stop_requested

bake.Run.stop_requested()

Whether stop() was asked (what the trainer checks at each step).

update

bake.Run.update(**fields)

Change run.json (the supervisor's; also stop).

wait

bake.Run.wait(poll=5.0, show=True)

Wait for the end; show progress; return the model.