Fine-Tuning, Distillation and Custom ModelsPreference and reinforcement fine-tuning · Lesson 10 of 16

Reinforcement fine-tuning with graders: RFT and GRPO

Article · 16 min · 8 min lecture

Video lecture

Reinforcement fine-tuning with graders: RFT and GRPO

14 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 14

Reinforcement fine-tuning

  • Checkable answers
  • GRPO and graders
  • Catching reward hacking

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

When the answer can be checked

Some tasks have verifiable outcomes: a calculation is right or wrong, code passes tests or fails, extracted fields match the source or do not, a reply cites the policy clause or does not. For these, you do not need humans to write or compare answers. You need a grader that scores outputs, and a training method that makes high-scoring outputs more likely. That is reinforcement fine-tuning (RFT). It is how many recent reasoning models were trained, and it is available to practitioners through open-source trainers and some hosted platforms.

GRPO in plain language

Group Relative Policy Optimization (GRPO) was introduced in the DeepSeekMath work (2024) and used in DeepSeek-R1's training. For each prompt:

  1. Sample a group of G completions from the current model (for example G = 8).
  2. Score each completion with your reward functions (graders).
  3. Compute each completion's advantage relative to its group: better than the group average → reinforce; worse → discourage.
  4. Update the model, optionally with a KL penalty to stay near a reference model.

Because advantages are computed within the group, GRPO needs no separate value (critic) model, which makes it simpler and lighter than PPO.

Designing graders (the whole game)

A grader is code (or a model) that returns a score. Good graders are:

  • Correct: they reward what you actually want. Test them on known good and bad outputs before training.
  • Hard to game: the model will find shortcuts. If you reward "contains the right number", it may list many numbers. Parse strictly.
  • Dense enough: partial credit (for example, per-field accuracy) often learns faster than all-or-nothing.
  • Composable: combine a correctness reward with format and length rewards, with weights.
TaskCorrectness graderGuard-rail graders
Fee quotesParsed amount equals fee-schedule calculationExactly one amount quoted; currency present
ExtractionPer-field match against goldValid JSON; no extra fields
SQL generationQuery runs on a sandbox DB and returns expected rowsRead-only statements only
Policy answersCites the correct clause IDLength under limit; no forbidden promises

Model-based graders (an LLM judge with a rubric) extend RFT to fuzzier tasks, but they are easier to game; calibrate them and prefer code where possible.

Hands-on: GRPO with TRL and custom rewards

{"prompt": [{"role": "system", "content": "Answer fee questions. Quote exactly one amount like 'PKR 25'."}, {"role": "user", "content": "Fee for sending PKR 12,000 to another bank?"}], "expected_fee": "PKR 25"}
# train_grpo.py  (pip install trl peft datasets transformers accelerate)
import os, re
from datasets import load_dataset
from peft import LoraConfig
from trl import GRPOConfig, GRPOTrainer

AMOUNT = re.compile(r"PKR\s?[\d,]+")

def text_of(completion):
    # conversational datasets give a list of messages per completion
    return completion[0]["content"] if isinstance(completion, list) else completion

def fee_correct(completions, expected_fee, **kwargs):
    rewards = []
    for c, exp in zip(completions, expected_fee):
        amounts = AMOUNT.findall(text_of(c))
        rewards.append(1.0 if len(amounts) == 1 and amounts[0].replace(" ", "") == exp.replace(" ", "") else 0.0)
    return rewards

def one_amount_and_short(completions, **kwargs):
    return [0.2 if len(AMOUNT.findall(text_of(c))) == 1 and len(text_of(c)) < 300 else 0.0 for c in completions]

data = load_dataset("json", data_files="fee_prompts.jsonl", split="train")   # values illustrative
args = GRPOConfig(output_dir="out/fees-grpo", num_generations=8, max_completion_length=128,
                  per_device_train_batch_size=8, gradient_accumulation_steps=4,
                  learning_rate=1e-5, num_train_epochs=1, logging_steps=5, bf16=True, report_to="none")
trainer = GRPOTrainer(model=os.getenv("SFT_MODEL", "out/fees-sft-merged"),
                      reward_funcs=[fee_correct, one_amount_and_short],
                      args=args, train_dataset=data,
                      peft_config=LoraConfig(r=16, lora_alpha=32, target_modules="all-linear", task_type="CAUSAL_LM"))
trainer.train()

Extra dataset columns (here expected_fee) are passed to reward functions as keyword arguments. Batch-size and generation settings must be compatible (TRL validates them); see the GRPOTrainer docs, and consider vLLM-accelerated generation for larger runs.

Hosted RFT options (check current availability)

  • Amazon Bedrock offers reinforcement fine-tuning for Amazon Nova models (starting with Nova 2 Lite per AWS documentation), with graders implemented as AWS Lambda functions or model-as-judge.
  • Microsoft Foundry (Azure) has offered reinforcement fine-tuning for selected reasoning models with gated access.
  • OpenAI offered RFT on selected reasoning models, but in 2026 began winding down its self-serve fine-tuning platform; do not plan new projects on it without confirming access.

Offerings change quickly; verify supported models, regions and pricing in current documentation.

Reward hacking: expect it

Symptoms: rewards climb but real quality does not; outputs become strange (repeating amounts, odd formats, excessive hedging). Countermeasures:

  • Inspect samples every few hundred steps, not just reward curves.
  • Keep an independent evaluation (different from the training grader) on held-out prompts.
  • Add guard-rail rewards and strict parsing; cap length.
  • Start from a good SFT model, so exploration begins from sensible behavior.

Worked example: extraction accuracy for a Karachi logistics firm

The firm extracts consignee, city, weight and COD amount from messy WhatsApp booking messages. SFT reached decent field accuracy but struggled with Urdu numerals and mixed units. They wrote a per-field grader (partial credit, strict parsing, unit normalization) and ran GRPO on 4,000 real messages. Field accuracy on a held-out set improved, with the largest gains on weights and COD amounts (illustrative). Early in training the model learned to output empty strings for uncertain fields (reward for "no wrong field"); they fixed the grader to penalize missing required fields.

Pitfalls

  • Using RFT where correctness cannot be checked reliably.
  • Trusting reward curves without reading samples.
  • Graders with loopholes (substring matches, lenient parsing).
  • Skipping SFT and starting RL from a model that rarely gets anything right (little signal).

How to measure success

An independent held-out evaluation shows gains over the SFT baseline on the target metric, guard-rail metrics (format, length, safety) are stable, and sample reviews show no reward hacking.

Key takeaways

  • RFT optimizes against graders when correctness is checkable
  • GRPO samples a group per prompt and reinforces above-average completions; no critic model needed
  • Graders must be correct, hard to game, dense and composable
  • Expect reward hacking: read samples and keep an independent evaluation
  • Hosted RFT exists on some platforms; availability changes, so verify

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. What makes GRPO lighter than PPO?
  2. During RFT, rewards rise steadily but reviewers see strange outputs listing many numbers. What is happening?
  3. Which task suits RFT best?

Put it into practice

Write a grader for one checkable task in your work, test it on 20 known good and bad outputs, and describe two ways a model could game it and how you would block them.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.