Fine-Tuning, Distillation and Custom ModelsPreference and reinforcement fine-tuning · Lesson 10 of 16
Reinforcement fine-tuning with graders: RFT and GRPO
Video lecture
Reinforcement fine-tuning with graders: RFT and GRPO
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Reinforcement fine-tuning
Some questions have a right answer you can check. Is the fee correct? Does the code pass its tests? Did the model extract the right weight from the booking? When you can check, you can train differently. Instead of writing ideal answers or comparing pairs, you write a grader, and let the model learn to score higher. That is reinforcement fine-tuning. In this lesson you will learn GRPO, how to design graders, how to run it with TRL, and how to catch reward hacking.
0:37 Analogy: eight attempts, one marker
Here is an analogy for GRPO. Imagine a class of students who each attempt the same maths question eight times. A strict marker scores every attempt. The teacher does not compare a student with the whole school, only with their own eight attempts: which were better than their own average? The student learns to do more of what worked. The marker is the grader, and if the marker only checks whether the right number appears anywhere, clever students will write every number they can think of.
1:14 GRPO in four steps
Here is GRPO, group relative policy optimization, introduced in DeepSeek's maths research and used in training DeepSeek R1. For each prompt, the model samples a group of answers, say eight. Your graders score each one. Each answer is compared with its own group's average. Better than average gets reinforced, worse gets discouraged. Then the model updates, optionally with a penalty to stay near a reference. Because the comparison is within the group, you do not need a separate critic model, which makes it simpler and lighter than PPO.
1:52 Good graders are
The grader is the whole game. A good grader is correct: it rewards what you actually want, so test it on known good and bad outputs first. It is hard to game: if you reward containing the right number, the model may list every number it can think of, so parse strictly. It is dense enough to learn from, so partial credit per field often beats all or nothing. And it is composable: correctness plus format plus length, with weights.
2:27 Grader examples
Examples. For fee quotes: the parsed amount must equal the fee schedule's calculation, with exactly one amount and a currency. For extraction: per-field match against gold, valid JSON and no extra fields. For SQL: the query runs on a sandbox database and returns the expected rows, and only read-only statements are allowed. For policy answers: it cites the right clause, stays under a length limit and makes no forbidden promises. A model can act as a judge for fuzzier tasks, but it is easier to game, so prefer code.
3:06 Hands-on: train_grpo.py
The TRL script has two reward functions. The first parses amounts with a strict pattern and gives one point only if exactly one amount is quoted and it matches the expected fee from the dataset. The second gives a small bonus for quoting one amount in a short answer. Extra dataset columns, like expected fee, arrive in your reward functions as keyword arguments. The config samples eight generations per prompt, caps completion length, and uses LoRA on top of your SFT model.
3:42 Hosted RFT (verify current docs)
Can you do this without running GPUs? Sometimes. Amazon Bedrock offers reinforcement fine-tuning for Amazon Nova models, starting with Nova 2 Lite according to AWS documentation, with graders as Lambda functions or a model judge. Microsoft's Foundry has offered it for selected reasoning models with gated access. OpenAI offered it too, but began winding down its self-serve fine-tuning platform in twenty twenty-six. These offerings change quickly, so always check current documentation.
4:13 Reward hacking countermeasures
Now expect reward hacking, because it will happen. Rewards climb, but real quality does not, and outputs get weird. Countermeasures: read samples every few hundred steps, not just the curves. Keep an independent evaluation, different from the training grader, on held-out prompts. Add guard-rail rewards, strict parsing and length caps. And always start from a decent SFT model, so exploration begins from sensible behavior.
4:41 Worked example: Karachi logistics (illustrative)
A worked example, with illustrative results. A Karachi logistics firm extracts consignee, city, weight and cash-on-delivery amount from messy WhatsApp bookings. SFT struggled with Urdu numerals and mixed units. A per-field grader with partial credit and unit normalization, plus GRPO on four thousand real messages, improved held-out accuracy, especially on weights and amounts. But early on the model learned to leave uncertain fields empty, since an empty field was never wrong. They fixed the grader to penalize missing required fields.
5:16 Simple example: percentage problems
A simple example. A tutoring app wants its small model to solve percentage problems like what is fifteen percent of two hundred forty. The grader is easy: compute the right answer in code and check the model's final number. They collect five hundred problems, run GRPO for an hour on top of an SFT model, and the pass rate on a separate test set rises. They also add a small reward for showing one clear final answer, which stops the model hedging between two numbers.
5:53 Run size and cost
How big does an RFT run need to be? Smaller than you might think for narrow tasks. A few hundred to a few thousand prompts can move a checkable metric, because each prompt produces a whole group of scored samples. The expensive part is generation: sampling eight answers per prompt costs far more than a single SFT pass. That is why fast generation, for example with vLLM integration, and short maximum completion lengths matter so much for cost.
6:27 Try this now
Try this now. Write your grader before you write any training code. Collect ten outputs you know are correct and ten you know are wrong, including one sneaky wrong answer that looks right. Run the grader on all twenty. If it scores the sneaky one as correct, fix the grader now, because training would find that loophole within hours.
6:53 Watch me do it
Watch me do it. Before any training, I write the fee grader and test it. Ten correct answers score one; ten wrong answers score zero; then I write a sneaky answer that lists three amounts including the right one. The grader gives it one. Loophole. I change it to require exactly one parsed amount, rerun, and the sneaky answer now scores zero. Next I start GRPO on our SFT model with eight generations per prompt and five hundred fee questions. Every hundred steps I stop to read twenty sampled answers, not just the reward curve. Around step three hundred, answers get shorter and one starts hedging with approximately; the length bonus is fine, but I add a small penalty for hedging words. At the end, I run the independent evaluation, a separate held-out set scored by a different grader, and pass rate is clearly above the SFT baseline. Samples look normal. Only then do I save the adapter.
8:02 Recap
Recap. Use reinforcement fine-tuning when correctness is checkable. GRPO reinforces answers that beat their group's average, without a critic model. Graders must be correct, hard to game, dense and composable. Expect reward hacking and keep an independent evaluation. Your next step: write a grader for one checkable task, test it on twenty good and bad outputs, and list two ways a model could game it.
When the answer can be checked
Some tasks have verifiable outcomes: a calculation is right or wrong, code passes tests or fails, extracted fields match the source or do not, a reply cites the policy clause or does not. For these, you do not need humans to write or compare answers. You need a grader that scores outputs, and a training method that makes high-scoring outputs more likely. That is reinforcement fine-tuning (RFT). It is how many recent reasoning models were trained, and it is available to practitioners through open-source trainers and some hosted platforms.
GRPO in plain language
Group Relative Policy Optimization (GRPO) was introduced in the DeepSeekMath work (2024) and used in DeepSeek-R1's training. For each prompt:
- Sample a group of G completions from the current model (for example G = 8).
- Score each completion with your reward functions (graders).
- Compute each completion's advantage relative to its group: better than the group average → reinforce; worse → discourage.
- Update the model, optionally with a KL penalty to stay near a reference model.
Because advantages are computed within the group, GRPO needs no separate value (critic) model, which makes it simpler and lighter than PPO.
Designing graders (the whole game)
A grader is code (or a model) that returns a score. Good graders are:
- Correct: they reward what you actually want. Test them on known good and bad outputs before training.
- Hard to game: the model will find shortcuts. If you reward "contains the right number", it may list many numbers. Parse strictly.
- Dense enough: partial credit (for example, per-field accuracy) often learns faster than all-or-nothing.
- Composable: combine a correctness reward with format and length rewards, with weights.
| Task | Correctness grader | Guard-rail graders |
|---|---|---|
| Fee quotes | Parsed amount equals fee-schedule calculation | Exactly one amount quoted; currency present |
| Extraction | Per-field match against gold | Valid JSON; no extra fields |
| SQL generation | Query runs on a sandbox DB and returns expected rows | Read-only statements only |
| Policy answers | Cites the correct clause ID | Length under limit; no forbidden promises |
Model-based graders (an LLM judge with a rubric) extend RFT to fuzzier tasks, but they are easier to game; calibrate them and prefer code where possible.
Hands-on: GRPO with TRL and custom rewards
{"prompt": [{"role": "system", "content": "Answer fee questions. Quote exactly one amount like 'PKR 25'."}, {"role": "user", "content": "Fee for sending PKR 12,000 to another bank?"}], "expected_fee": "PKR 25"}# train_grpo.py (pip install trl peft datasets transformers accelerate)
import os, re
from datasets import load_dataset
from peft import LoraConfig
from trl import GRPOConfig, GRPOTrainer
AMOUNT = re.compile(r"PKR\s?[\d,]+")
def text_of(completion):
# conversational datasets give a list of messages per completion
return completion[0]["content"] if isinstance(completion, list) else completion
def fee_correct(completions, expected_fee, **kwargs):
rewards = []
for c, exp in zip(completions, expected_fee):
amounts = AMOUNT.findall(text_of(c))
rewards.append(1.0 if len(amounts) == 1 and amounts[0].replace(" ", "") == exp.replace(" ", "") else 0.0)
return rewards
def one_amount_and_short(completions, **kwargs):
return [0.2 if len(AMOUNT.findall(text_of(c))) == 1 and len(text_of(c)) < 300 else 0.0 for c in completions]
data = load_dataset("json", data_files="fee_prompts.jsonl", split="train") # values illustrative
args = GRPOConfig(output_dir="out/fees-grpo", num_generations=8, max_completion_length=128,
per_device_train_batch_size=8, gradient_accumulation_steps=4,
learning_rate=1e-5, num_train_epochs=1, logging_steps=5, bf16=True, report_to="none")
trainer = GRPOTrainer(model=os.getenv("SFT_MODEL", "out/fees-sft-merged"),
reward_funcs=[fee_correct, one_amount_and_short],
args=args, train_dataset=data,
peft_config=LoraConfig(r=16, lora_alpha=32, target_modules="all-linear", task_type="CAUSAL_LM"))
trainer.train()Extra dataset columns (here expected_fee) are passed to reward functions as keyword arguments. Batch-size and generation settings must be compatible (TRL validates them); see the GRPOTrainer docs, and consider vLLM-accelerated generation for larger runs.
Hosted RFT options (check current availability)
- Amazon Bedrock offers reinforcement fine-tuning for Amazon Nova models (starting with Nova 2 Lite per AWS documentation), with graders implemented as AWS Lambda functions or model-as-judge.
- Microsoft Foundry (Azure) has offered reinforcement fine-tuning for selected reasoning models with gated access.
- OpenAI offered RFT on selected reasoning models, but in 2026 began winding down its self-serve fine-tuning platform; do not plan new projects on it without confirming access.
Offerings change quickly; verify supported models, regions and pricing in current documentation.
Reward hacking: expect it
Symptoms: rewards climb but real quality does not; outputs become strange (repeating amounts, odd formats, excessive hedging). Countermeasures:
- Inspect samples every few hundred steps, not just reward curves.
- Keep an independent evaluation (different from the training grader) on held-out prompts.
- Add guard-rail rewards and strict parsing; cap length.
- Start from a good SFT model, so exploration begins from sensible behavior.
Worked example: extraction accuracy for a Karachi logistics firm
The firm extracts consignee, city, weight and COD amount from messy WhatsApp booking messages. SFT reached decent field accuracy but struggled with Urdu numerals and mixed units. They wrote a per-field grader (partial credit, strict parsing, unit normalization) and ran GRPO on 4,000 real messages. Field accuracy on a held-out set improved, with the largest gains on weights and COD amounts (illustrative). Early in training the model learned to output empty strings for uncertain fields (reward for "no wrong field"); they fixed the grader to penalize missing required fields.
Pitfalls
- Using RFT where correctness cannot be checked reliably.
- Trusting reward curves without reading samples.
- Graders with loopholes (substring matches, lenient parsing).
- Skipping SFT and starting RL from a model that rarely gets anything right (little signal).
How to measure success
An independent held-out evaluation shows gains over the SFT baseline on the target metric, guard-rail metrics (format, length, safety) are stable, and sample reviews show no reward hacking.
Key takeaways
- RFT optimizes against graders when correctness is checkable
- GRPO samples a group per prompt and reinforces above-average completions; no critic model needed
- Graders must be correct, hard to game, dense and composable
- Expect reward hacking: read samples and keep an independent evaluation
- Hosted RFT exists on some platforms; availability changes, so verify
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Write a grader for one checkable task in your work, test it on 20 known good and bad outputs, and describe two ways a model could game it and how you would block them.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.