Skip to content

Optimizing a prompt: English to Québécois French

English to Québécois French: an AI judge as the metric, InstructionSearch, before and after.

A small model translates English into standard French. We want everyday Québécois French (“dépanneur”, “frette”, “char”), and we have twenty examples of it. An AI judge scores translations, and an optimizer searches for the instruction that raises the score. Every step is measured, with its uncertainty.

Every output below is a real reply. This page is a notebook: open it in Chattering and run it, or run it all with python/.venv/bin/python tools/docs.py run python/examples/optimizing_translator/README.md.

import functai
functai.configure(lm="gpt-4.1-mini", temperature=0)

from functai import ai, _ai

@ai(lm="gpt-4.1-nano")
def translator(english: str) -> str:
    """Translate to French."""
    ...

The examples

One dict per row: english is the translator’s input, result the translation we want.

rows = [
    {"english": "I'm going to the convenience store.", "result": "Je m'en vais au dépanneur."},
    {"english": "It's really cold out today.", "result": "Il fait frette en maudit aujourd'hui."},
    {"english": "Can you help me move this weekend?", "result": "Tu peux m'aider à déménager ce weekend?"},
    {"english": "We were stuck in traffic for two hours.", "result": "On était pognés dans le trafic pendant deux heures."},
    {"english": "She's my girlfriend.", "result": "C'est ma blonde."},
    {"english": "That car is so cool!", "result": "C'est ben l'fun ce char-là!"},
    {"english": "I'll call you tonight.", "result": "Je vais t'appeler ce soir."},
    {"english": "He's always bragging.", "result": "Il se vente tout l'temps."},
    {"english": "We grabbed a coffee at Tim's.", "result": "On a pris un café au Tim."},
    {"english": "Close the window, it's chilly.", "result": "Ferme la fenêtre, y fait frette."},
    {"english": "I have an appointment at 3.", "result": "J'ai un rendez-vous à trois heures."},
    {"english": "They're celebrating their birthday.", "result": "Ils fêtent leur fête."},
    {"english": "I parked in the back.", "result": "J'ai stationné dans l'fond."},
    {"english": "The metro is packed.", "result": "Le métro est plein à craquer."},
    {"english": "We watched a movie last night.", "result": "On a écouté un film hier soir."},
    {"english": "I need to do my groceries.", "result": "J'dois faire mon épicerie."},
    {"english": "Don't forget your boots.", "result": "Oublie pas tes bottes."},
    {"english": "It's snowing again.", "result": "Il neige encore."},
    {"english": "I'll take the bus.", "result": "J'va prendre l'bus."},
    {"english": "We're out of milk.", "result": "On est à court de lait."},
]
train, dev = rows[:10], rows[10:]

The judge

Exact match is too strict for translations, so the metric is an AI function too. It reads the row (with the expected result) and the prediction:

@ai
def judge(row: dict, prediction: dict) -> float:
    """How well the prediction matches row['result']: the same meaning,
    in the same everyday Québécois register. 1 is as good, 0 is wrong or
    in standard French."""
    score: float = _ai["Between 0 and 1."]
    return max(0.0, min(1.0, score))

Before

before = functai.evaluate(translator, dev, judge, num_threads=5)
before
Evaluation(translator, 10 examples: judge 0.38 [0.03, 0.73])
from dpyr import col

before.table.arrange(col.judge).select(col.english, col.result, col.pred_result, col.judge)
# dpyr dataframe · source: polars · showing 10 of 10 rows
┌─────────────────────────────────────┬─────────────────────────────────────┬───────────────────────────────────────┬───────┐
│ english                             ┆ result                              ┆ pred_result                           ┆ judge │
│ ---                                 ┆ ---                                 ┆ ---                                   ┆ ---   │
│ str                                 ┆ str                                 ┆ str                                   ┆ f64   │
╞═════════════════════════════════════╪═════════════════════════════════════╪═══════════════════════════════════════╪═══════╡
│ I have an appointment at 3.         ┆ J'ai un rendez-vous à trois heures. ┆ J'ai un rendez-vous à 15 heures.      ┆ 0.0   │
│ I parked in the back.               ┆ J'ai stationné dans l'fond.         ┆ Je me suis garé à l'arrière.          ┆ 0.0   │
│ We watched a movie last night.      ┆ On a écouté un film hier soir.      ┆ Nous avons regardé un film hier soir. ┆ 0.0   │
│ I need to do my groceries.          ┆ J'dois faire mon épicerie.          ┆ Je dois faire mes courses.            ┆ 0.0   │
│ Don't forget your boots.            ┆ Oublie pas tes bottes.              ┆ N'oublie pas tes bottes.              ┆ 0.0   │
│ I'll take the bus.                  ┆ J'va prendre l'bus.                 ┆ Je vais prendre le bus.               ┆ 0.0   │
│ We're out of milk.                  ┆ On est à court de lait.             ┆ Nous n'avons plus de lait.            ┆ 0.8   │
│ They're celebrating their birthday. ┆ Ils fêtent leur fête.               ┆ Ils fêtent leur anniversaire.         ┆ 1.0   │
│ The metro is packed.                ┆ Le métro est plein à craquer.       ┆ Le métro est bondé.                   ┆ 1.0   │
│ It's snowing again.                 ┆ Il neige encore.                    ┆ Il neige encore.                      ┆ 1.0   │
└─────────────────────────────────────┴─────────────────────────────────────┴───────────────────────────────────────┴───────┘

Searching for a better instruction

InstructionSearch asks a model (prompt_lm) to propose instructions from the function’s code and a few examples, then tries them on minibatches of the data and keeps the best. Setting both demo limits to 0 means only the instruction changes.

from functai import InstructionSearch

opt = InstructionSearch(num_candidates=4, num_trials=8, minibatch_size=5,
                        max_bootstrapped_demos=0, max_labeled_demos=0,
                        prompt_lm="gpt-4.1-mini")
better = translator.opt(train, metric=judge, optimizer=opt)
print(better.instructions)
Function: translator

Translate the given English sentence into informal Quebec French (joual), capturing local slang, expressions, and conversational style for a natural, idiomatic translation.

Every trial is a row (combo says which instruction and which demo set):

from dpyr import read

read(better.trials)
# dpyr dataframe · source: polars · showing 8 of 8 rows
┌───────┬───────────┬─────────────────┬────────────┐
│ trial ┆ combo     ┆ minibatch_score ┆ full_score │
│ ---   ┆ ---       ┆ ---             ┆ ---        │
│ i64   ┆ list[i64] ┆ f64             ┆ f64        │
╞═══════╪═══════════╪═════════════════╪════════════╡
│ 0     ┆ [0, 0]    ┆ 0.2             ┆ null       │
│ 1     ┆ [1, 0]    ┆ 1.0             ┆ 0.96       │
│ 2     ┆ [0, 0]    ┆ 0.2             ┆ null       │
│ 3     ┆ [1, 0]    ┆ 0.96            ┆ 0.96       │
│ 4     ┆ [3, 0]    ┆ 0.96            ┆ 0.86       │
│ 5     ┆ [2, 0]    ┆ 0.92            ┆ 0.94       │
│ 6     ┆ [1, 0]    ┆ 0.92            ┆ 0.96       │
│ 7     ┆ [0, 0]    ┆ 0.0             ┆ null       │
└───────┴───────────┴─────────────────┴────────────┘

After

after = functai.evaluate(better, dev, judge, num_threads=5)
functai.compare(before, after)
# dpyr dataframe · source: polars · showing 1 of 1 rows
┌────────┬────────┬───────┬──────┬──────────┬──────────┬────────┬───────┬──────┬─────┐
│ metric ┆ before ┆ after ┆ diff ┆ low      ┆ high     ┆ better ┆ worse ┆ same ┆ n   │
│ ---    ┆ ---    ┆ ---   ┆ ---  ┆ ---      ┆ ---      ┆ ---    ┆ ---   ┆ ---  ┆ --- │
│ str    ┆ f64    ┆ f64   ┆ f64  ┆ f64      ┆ f64      ┆ i64    ┆ i64   ┆ i64  ┆ i64 │
╞════════╪════════╪═══════╪══════╪══════════╪══════════╪════════╪═══════╪══════╪═════╡
│ judge  ┆ 0.38   ┆ 0.86  ┆ 0.48 ┆ 0.109388 ┆ 0.850612 ┆ 6      ┆ 1     ┆ 3    ┆ 10  │
└────────┴────────┴───────┴──────┴──────────┴──────────┴────────┴───────┴──────┴─────┘

diff is the change in the judge’s mean over the same ten examples, with its 95% interval; better, worse and same count examples.

better("Hi, what's the weather like? I'm going to the convenience store.")
"Salut, c'est quoi le temps qu'il fait? Je m'en vais au dépanneur."

translator still has the old instruction; better.save("translator.json") keeps the new one.