---
title: "Reinforcement fine-tuning with graders: RFT and GRPO"
description: "When the answer can be checked Some tasks have verifiable outcomes: a calculation is right or wrong, code passes tests or fails, extracted fields match…"
url: https://optimizeall.com/learn/fine-tuning-and-custom-models/rft-and-grpo
updated: 2026-10-05
---

Fine-Tuning, Distillation and Custom Models · Preference and reinforcement fine-tuning · lesson 10 of 16 · 16 min

# Reinforcement fine-tuning with graders: RFT and GRPO

## When the answer can be checked

Some tasks have **verifiable** outcomes: a calculation is right or wrong, code passes tests or fails, extracted fields match the source or do not, a reply cites the policy clause or does not. For these, you do not need humans to write or compare answers. You need a **grader** that scores outputs, and a training method that makes high-scoring outputs more likely. That is **reinforcement fine-tuning (RFT)**. It is how many recent reasoning models were trained, and it is available to practitioners through open-source trainers and some hosted platforms.

## GRPO in plain language

**Group Relative Policy Optimization (GRPO)** was introduced in the DeepSeekMath work (2024) and used in DeepSeek-R1's training. For each prompt:

1. Sample a **group** of G completions from the current model (for example G = 8).
2. Score each completion with your **reward functions** (graders).
3. Compute each completion's **advantage** relative to its group: better than the group average → reinforce; worse → discourage.
4. Update the model, optionally with a KL penalty to stay near a reference model.

Because advantages are computed within the group, GRPO needs **no separate value (critic) model**, which makes it simpler and lighter than PPO.

## Designing graders (the whole game)

A grader is code (or a model) that returns a score. Good graders are:

- **Correct:** they reward what you actually want. Test them on known good and bad outputs before training.
- **Hard to game:** the model will find shortcuts. If you reward "contains the right number", it may list many numbers. Parse strictly.
- **Dense enough:** partial credit (for example, per-field accuracy) often learns faster than all-or-nothing.
- **Composable:** combine a correctness reward with format and length rewards, with weights.

| Task | Correctness grader | Guard-rail graders |
|---|---|---|
| Fee quotes | Parsed amount equals fee-schedule calculation | Exactly one amount quoted; currency present |
| Extraction | Per-field match against gold | Valid JSON; no extra fields |
| SQL generation | Query runs on a sandbox DB and returns expected rows | Read-only statements only |
| Policy answers | Cites the correct clause ID | Length under limit; no forbidden promises |

Model-based graders (an LLM judge with a rubric) extend RFT to fuzzier tasks, but they are easier to game; calibrate them and prefer code where possible.

## Hands-on: GRPO with TRL and custom rewards

```jsonl
{"prompt": [{"role": "system", "content": "Answer fee questions. Quote exactly one amount like 'PKR 25'."}, {"role": "user", "content": "Fee for sending PKR 12,000 to another bank?"}], "expected_fee": "PKR 25"}
```

```python
# train_grpo.py  (pip install trl peft datasets transformers accelerate)
import os, re
from datasets import load_dataset
from peft import LoraConfig
from trl import GRPOConfig, GRPOTrainer

AMOUNT = re.compile(r"PKR\s?[\d,]+")

def text_of(completion):
    # conversational datasets give a list of messages per completion
    return completion[0]["content"] if isinstance(completion, list) else completion

def fee_correct(completions, expected_fee, **kwargs):
    rewards = []
    for c, exp in zip(completions, expected_fee):
        amounts = AMOUNT.findall(text_of(c))
        rewards.append(1.0 if len(amounts) == 1 and amounts[0].replace(" ", "") == exp.replace(" ", "") else 0.0)
    return rewards

def one_amount_and_short(completions, **kwargs):
    return [0.2 if len(AMOUNT.findall(text_of(c))) == 1 and len(text_of(c)) < 300 else 0.0 for c in completions]

data = load_dataset("json", data_files="fee_prompts.jsonl", split="train")   # values illustrative
args = GRPOConfig(output_dir="out/fees-grpo", num_generations=8, max_completion_length=128,
                  per_device_train_batch_size=8, gradient_accumulation_steps=4,
                  learning_rate=1e-5, num_train_epochs=1, logging_steps=5, bf16=True, report_to="none")
trainer = GRPOTrainer(model=os.getenv("SFT_MODEL", "out/fees-sft-merged"),
                      reward_funcs=[fee_correct, one_amount_and_short],
                      args=args, train_dataset=data,
                      peft_config=LoraConfig(r=16, lora_alpha=32, target_modules="all-linear", task_type="CAUSAL_LM"))
trainer.train()
```

Extra dataset columns (here `expected_fee`) are passed to reward functions as keyword arguments. Batch-size and generation settings must be compatible (TRL validates them); see the GRPOTrainer docs, and consider vLLM-accelerated generation for larger runs.

## Hosted RFT options (check current availability)

- **Amazon Bedrock** offers reinforcement fine-tuning for Amazon Nova models (starting with Nova 2 Lite per AWS documentation), with graders implemented as AWS Lambda functions or model-as-judge.
- **Microsoft Foundry** (Azure) has offered reinforcement fine-tuning for selected reasoning models with gated access.
- **OpenAI** offered RFT on selected reasoning models, but in 2026 began winding down its self-serve fine-tuning platform; do not plan new projects on it without confirming access.

Offerings change quickly; verify supported models, regions and pricing in current documentation.

## Reward hacking: expect it

Symptoms: rewards climb but real quality does not; outputs become strange (repeating amounts, odd formats, excessive hedging). Countermeasures:

- Inspect samples every few hundred steps, not just reward curves.
- Keep an **independent evaluation** (different from the training grader) on held-out prompts.
- Add guard-rail rewards and strict parsing; cap length.
- Start from a good SFT model, so exploration begins from sensible behavior.

## Worked example: extraction accuracy for a Karachi logistics firm

The firm extracts consignee, city, weight and COD amount from messy WhatsApp booking messages. SFT reached decent field accuracy but struggled with Urdu numerals and mixed units. They wrote a per-field grader (partial credit, strict parsing, unit normalization) and ran GRPO on 4,000 real messages. Field accuracy on a held-out set improved, with the largest gains on weights and COD amounts (illustrative). Early in training the model learned to output empty strings for uncertain fields (reward for "no wrong field"); they fixed the grader to penalize missing required fields.

## Pitfalls

- Using RFT where correctness cannot be checked reliably.
- Trusting reward curves without reading samples.
- Graders with loopholes (substring matches, lenient parsing).
- Skipping SFT and starting RL from a model that rarely gets anything right (little signal).

## How to measure success

An independent held-out evaluation shows gains over the SFT baseline on the target metric, guard-rail metrics (format, length, safety) are stable, and sample reviews show no reward hacking.

## Video lecture: Reinforcement fine-tuning with graders: RFT and GRPO

Lecture coming soon · 14 chapters · about 8 minutes. Read the full transcript below.

1. Reinforcement fine-tuning
2. Analogy: eight attempts, one marker
3. GRPO in four steps
4. Good graders are
5. Grader examples
6. Hands-on: train_grpo.py
7. Hosted RFT (verify current docs)
8. Reward hacking countermeasures
9. Worked example: Karachi logistics (illustrative)
10. Simple example: percentage problems
11. Run size and cost
12. Try this now
13. Watch me do it
14. Recap

## Lecture transcript

### Reinforcement fine-tuning

Some questions have a right answer you can check. Is the fee correct? Does the code pass its tests? Did the model extract the right weight from the booking? When you can check, you can train differently. Instead of writing ideal answers or comparing pairs, you write a grader, and let the model learn to score higher. That is reinforcement fine-tuning. In this lesson you will learn GRPO, how to design graders, how to run it with TRL, and how to catch reward hacking.

### Analogy: eight attempts, one marker

Here is an analogy for GRPO. Imagine a class of students who each attempt the same maths question eight times. A strict marker scores every attempt. The teacher does not compare a student with the whole school, only with their own eight attempts: which were better than their own average? The student learns to do more of what worked. The marker is the grader, and if the marker only checks whether the right number appears anywhere, clever students will write every number they can think of.

### GRPO in four steps

Here is GRPO, group relative policy optimization, introduced in DeepSeek's maths research and used in training DeepSeek R1. For each prompt, the model samples a group of answers, say eight. Your graders score each one. Each answer is compared with its own group's average. Better than average gets reinforced, worse gets discouraged. Then the model updates, optionally with a penalty to stay near a reference. Because the comparison is within the group, you do not need a separate critic model, which makes it simpler and lighter than PPO.

### Good graders are

The grader is the whole game. A good grader is correct: it rewards what you actually want, so test it on known good and bad outputs first. It is hard to game: if you reward containing the right number, the model may list every number it can think of, so parse strictly. It is dense enough to learn from, so partial credit per field often beats all or nothing. And it is composable: correctness plus format plus length, with weights.

### Grader examples

Examples. For fee quotes: the parsed amount must equal the fee schedule's calculation, with exactly one amount and a currency. For extraction: per-field match against gold, valid JSON and no extra fields. For SQL: the query runs on a sandbox database and returns the expected rows, and only read-only statements are allowed. For policy answers: it cites the right clause, stays under a length limit and makes no forbidden promises. A model can act as a judge for fuzzier tasks, but it is easier to game, so prefer code.

### Hands-on: train_grpo.py

The TRL script has two reward functions. The first parses amounts with a strict pattern and gives one point only if exactly one amount is quoted and it matches the expected fee from the dataset. The second gives a small bonus for quoting one amount in a short answer. Extra dataset columns, like expected fee, arrive in your reward functions as keyword arguments. The config samples eight generations per prompt, caps completion length, and uses LoRA on top of your SFT model.

### Hosted RFT (verify current docs)

Can you do this without running GPUs? Sometimes. Amazon Bedrock offers reinforcement fine-tuning for Amazon Nova models, starting with Nova 2 Lite according to AWS documentation, with graders as Lambda functions or a model judge. Microsoft's Foundry has offered it for selected reasoning models with gated access. OpenAI offered it too, but began winding down its self-serve fine-tuning platform in twenty twenty-six. These offerings change quickly, so always check current documentation.

### Reward hacking countermeasures

Now expect reward hacking, because it will happen. Rewards climb, but real quality does not, and outputs get weird. Countermeasures: read samples every few hundred steps, not just the curves. Keep an independent evaluation, different from the training grader, on held-out prompts. Add guard-rail rewards, strict parsing and length caps. And always start from a decent SFT model, so exploration begins from sensible behavior.

### Worked example: Karachi logistics (illustrative)

A worked example, with illustrative results. A Karachi logistics firm extracts consignee, city, weight and cash-on-delivery amount from messy WhatsApp bookings. SFT struggled with Urdu numerals and mixed units. A per-field grader with partial credit and unit normalization, plus GRPO on four thousand real messages, improved held-out accuracy, especially on weights and amounts. But early on the model learned to leave uncertain fields empty, since an empty field was never wrong. They fixed the grader to penalize missing required fields.

### Simple example: percentage problems

A simple example. A tutoring app wants its small model to solve percentage problems like what is fifteen percent of two hundred forty. The grader is easy: compute the right answer in code and check the model's final number. They collect five hundred problems, run GRPO for an hour on top of an SFT model, and the pass rate on a separate test set rises. They also add a small reward for showing one clear final answer, which stops the model hedging between two numbers.

### Run size and cost

How big does an RFT run need to be? Smaller than you might think for narrow tasks. A few hundred to a few thousand prompts can move a checkable metric, because each prompt produces a whole group of scored samples. The expensive part is generation: sampling eight answers per prompt costs far more than a single SFT pass. That is why fast generation, for example with vLLM integration, and short maximum completion lengths matter so much for cost.

### Try this now

Try this now. Write your grader before you write any training code. Collect ten outputs you know are correct and ten you know are wrong, including one sneaky wrong answer that looks right. Run the grader on all twenty. If it scores the sneaky one as correct, fix the grader now, because training would find that loophole within hours.

### Watch me do it

Watch me do it. Before any training, I write the fee grader and test it. Ten correct answers score one; ten wrong answers score zero; then I write a sneaky answer that lists three amounts including the right one. The grader gives it one. Loophole. I change it to require exactly one parsed amount, rerun, and the sneaky answer now scores zero. Next I start GRPO on our SFT model with eight generations per prompt and five hundred fee questions. Every hundred steps I stop to read twenty sampled answers, not just the reward curve. Around step three hundred, answers get shorter and one starts hedging with approximately; the length bonus is fine, but I add a small penalty for hedging words. At the end, I run the independent evaluation, a separate held-out set scored by a different grader, and pass rate is clearly above the SFT baseline. Samples look normal. Only then do I save the adapter.

### Recap

Recap. Use reinforcement fine-tuning when correctness is checkable. GRPO reinforces answers that beat their group's average, without a critic model. Graders must be correct, hard to game, dense and composable. Expect reward hacking and keep an independent evaluation. Your next step: write a grader for one checkable task, test it on twenty good and bad outputs, and list two ways a model could game it.

## Key takeaways

- RFT optimizes against graders when correctness is checkable
- GRPO samples a group per prompt and reinforces above-average completions; no critic model needed
- Graders must be correct, hard to game, dense and composable
- Expect reward hacking: read samples and keep an independent evaluation
- Hosted RFT exists on some platforms; availability changes, so verify

## Try it

Write a grader for one checkable task in your work, test it on 20 known good and bad outputs, and describe two ways a model could game it and how you would block them.

- [Previous: Preference optimization: RLHF and DPO for practitioners](https://optimizeall.com/learn/fine-tuning-and-custom-models/rlhf-and-dpo-explained)
- [Next: Evaluation before and after fine-tuning](https://optimizeall.com/learn/fine-tuning-and-custom-models/evaluation-before-and-after)
- [All lessons of Fine-Tuning, Distillation and Custom Models](https://optimizeall.com/learn/fine-tuning-and-custom-models)
