Optimizing a prompt: English to Québécois French¶
English to Québécois French: an AI judge as the metric, InstructionSearch, before and after.
A small model translates English into standard French. We want everyday Québécois French (“dépanneur”, “frette”, “char”), and we have twenty examples of it. An AI judge scores translations, and an optimizer searches for the instruction that raises the score. Every step is measured, with its uncertainty.
Every output below is a real reply. This page is a notebook: open it in
Chattering and run it, or run it all with python/.venv/bin/python tools/docs.py run python/examples/optimizing_translator/README.md.
import functai
functai.configure(lm="gpt-4.1-mini", temperature=0)
from functai import ai, _ai
@ai(lm="gpt-4.1-nano")
def translator(english: str) -> str:
"""Translate to French."""
...
The examples¶
One dict per row: english is the translator’s input, result the
translation we want.
rows = [
{"english": "I'm going to the convenience store.", "result": "Je m'en vais au dépanneur."},
{"english": "It's really cold out today.", "result": "Il fait frette en maudit aujourd'hui."},
{"english": "Can you help me move this weekend?", "result": "Tu peux m'aider à déménager ce weekend?"},
{"english": "We were stuck in traffic for two hours.", "result": "On était pognés dans le trafic pendant deux heures."},
{"english": "She's my girlfriend.", "result": "C'est ma blonde."},
{"english": "That car is so cool!", "result": "C'est ben l'fun ce char-là!"},
{"english": "I'll call you tonight.", "result": "Je vais t'appeler ce soir."},
{"english": "He's always bragging.", "result": "Il se vente tout l'temps."},
{"english": "We grabbed a coffee at Tim's.", "result": "On a pris un café au Tim."},
{"english": "Close the window, it's chilly.", "result": "Ferme la fenêtre, y fait frette."},
{"english": "I have an appointment at 3.", "result": "J'ai un rendez-vous à trois heures."},
{"english": "They're celebrating their birthday.", "result": "Ils fêtent leur fête."},
{"english": "I parked in the back.", "result": "J'ai stationné dans l'fond."},
{"english": "The metro is packed.", "result": "Le métro est plein à craquer."},
{"english": "We watched a movie last night.", "result": "On a écouté un film hier soir."},
{"english": "I need to do my groceries.", "result": "J'dois faire mon épicerie."},
{"english": "Don't forget your boots.", "result": "Oublie pas tes bottes."},
{"english": "It's snowing again.", "result": "Il neige encore."},
{"english": "I'll take the bus.", "result": "J'va prendre l'bus."},
{"english": "We're out of milk.", "result": "On est à court de lait."},
]
train, dev = rows[:10], rows[10:]
The judge¶
Exact match is too strict for translations, so the metric is an AI
function too. It reads the row (with the expected result) and the
prediction:
@ai
def judge(row: dict, prediction: dict) -> float:
"""How well the prediction matches row['result']: the same meaning,
in the same everyday Québécois register. 1 is as good, 0 is wrong or
in standard French."""
score: float = _ai["Between 0 and 1."]
return max(0.0, min(1.0, score))
Before¶
before = functai.evaluate(translator, dev, judge, num_threads=5)
before
Evaluation(translator, 10 examples: judge 0.38 [0.03, 0.73])
from dpyr import col
before.table.arrange(col.judge).select(col.english, col.result, col.pred_result, col.judge)
# dpyr dataframe · source: polars · showing 10 of 10 rows
┌─────────────────────────────────────┬─────────────────────────────────────┬───────────────────────────────────────┬───────┐
│ english ┆ result ┆ pred_result ┆ judge │
│ --- ┆ --- ┆ --- ┆ --- │
│ str ┆ str ┆ str ┆ f64 │
╞═════════════════════════════════════╪═════════════════════════════════════╪═══════════════════════════════════════╪═══════╡
│ I have an appointment at 3. ┆ J'ai un rendez-vous à trois heures. ┆ J'ai un rendez-vous à 15 heures. ┆ 0.0 │
│ I parked in the back. ┆ J'ai stationné dans l'fond. ┆ Je me suis garé à l'arrière. ┆ 0.0 │
│ We watched a movie last night. ┆ On a écouté un film hier soir. ┆ Nous avons regardé un film hier soir. ┆ 0.0 │
│ I need to do my groceries. ┆ J'dois faire mon épicerie. ┆ Je dois faire mes courses. ┆ 0.0 │
│ Don't forget your boots. ┆ Oublie pas tes bottes. ┆ N'oublie pas tes bottes. ┆ 0.0 │
│ I'll take the bus. ┆ J'va prendre l'bus. ┆ Je vais prendre le bus. ┆ 0.0 │
│ We're out of milk. ┆ On est à court de lait. ┆ Nous n'avons plus de lait. ┆ 0.8 │
│ They're celebrating their birthday. ┆ Ils fêtent leur fête. ┆ Ils fêtent leur anniversaire. ┆ 1.0 │
│ The metro is packed. ┆ Le métro est plein à craquer. ┆ Le métro est bondé. ┆ 1.0 │
│ It's snowing again. ┆ Il neige encore. ┆ Il neige encore. ┆ 1.0 │
└─────────────────────────────────────┴─────────────────────────────────────┴───────────────────────────────────────┴───────┘
Searching for a better instruction¶
InstructionSearch asks a model (prompt_lm) to propose instructions
from the function’s code and a few examples, then tries them on
minibatches of the data and keeps the best. Setting both demo limits to
0 means only the instruction changes.
from functai import InstructionSearch
opt = InstructionSearch(num_candidates=4, num_trials=8, minibatch_size=5,
max_bootstrapped_demos=0, max_labeled_demos=0,
prompt_lm="gpt-4.1-mini")
better = translator.opt(train, metric=judge, optimizer=opt)
print(better.instructions)
Function: translator
Translate the given English sentence into informal Quebec French (joual), capturing local slang, expressions, and conversational style for a natural, idiomatic translation.
Every trial is a row (combo says which instruction and which demo
set):
from dpyr import read
read(better.trials)
# dpyr dataframe · source: polars · showing 8 of 8 rows
┌───────┬───────────┬─────────────────┬────────────┐
│ trial ┆ combo ┆ minibatch_score ┆ full_score │
│ --- ┆ --- ┆ --- ┆ --- │
│ i64 ┆ list[i64] ┆ f64 ┆ f64 │
╞═══════╪═══════════╪═════════════════╪════════════╡
│ 0 ┆ [0, 0] ┆ 0.2 ┆ null │
│ 1 ┆ [1, 0] ┆ 1.0 ┆ 0.96 │
│ 2 ┆ [0, 0] ┆ 0.2 ┆ null │
│ 3 ┆ [1, 0] ┆ 0.96 ┆ 0.96 │
│ 4 ┆ [3, 0] ┆ 0.96 ┆ 0.86 │
│ 5 ┆ [2, 0] ┆ 0.92 ┆ 0.94 │
│ 6 ┆ [1, 0] ┆ 0.92 ┆ 0.96 │
│ 7 ┆ [0, 0] ┆ 0.0 ┆ null │
└───────┴───────────┴─────────────────┴────────────┘
After¶
after = functai.evaluate(better, dev, judge, num_threads=5)
functai.compare(before, after)
# dpyr dataframe · source: polars · showing 1 of 1 rows
┌────────┬────────┬───────┬──────┬──────────┬──────────┬────────┬───────┬──────┬─────┐
│ metric ┆ before ┆ after ┆ diff ┆ low ┆ high ┆ better ┆ worse ┆ same ┆ n │
│ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- │
│ str ┆ f64 ┆ f64 ┆ f64 ┆ f64 ┆ f64 ┆ i64 ┆ i64 ┆ i64 ┆ i64 │
╞════════╪════════╪═══════╪══════╪══════════╪══════════╪════════╪═══════╪══════╪═════╡
│ judge ┆ 0.38 ┆ 0.86 ┆ 0.48 ┆ 0.109388 ┆ 0.850612 ┆ 6 ┆ 1 ┆ 3 ┆ 10 │
└────────┴────────┴───────┴──────┴──────────┴──────────┴────────┴───────┴──────┴─────┘
diff is the change in the judge’s mean over the same ten examples,
with its 95% interval; better, worse and same count examples.
better("Hi, what's the weather like? I'm going to the convenience store.")
"Salut, c'est quoi le temps qu'il fait? Je m'en vais au dépanneur."
translator still has the old instruction;
better.save("translator.json") keeps the new one.