Last month we released the Document Agent, a general agent that can perform multi-step document-related tasks. Under the hood, an LLM works in our harness step by step: it parses pages, writes scripts, and runs checks until the output passes.
As an experiment, we fine-tuned a smaller open model to drive the agent. We tried two approaches: training on a stronger model’s runs, and reinforcement learning. Both helped, but only RL consistently stopped the model from getting stuck repeating the same tool calls.
Setup
For this experiment we trained on one of our agent tasks: converting scientific PDFs to JATS XML.
Out of the box, the smaller model isn’t very good at the task at hand, mostly because it’s just not that competent with the tools it’s given. It gets stuck, reruns the same tools, and, most nefariously, gets caught in a loop of calling the same tool over and over again. Looping is a fairly common issue in LLMs; see Liquid AI’s post on doom loops.
There are two main ways to improve the model driving the agent:
- Learn from a stronger model (supervised fine-tuning, SFT). We record a stronger model doing the task and fine-tune the small model to copy it, step by step.
- Learn from its own mistakes (reinforcement learning, RL). The small model attempts the task many times, each attempt is scored, and the model is pushed toward the attempts that scored well.
We tried SFT, RL from the base model, and SFT followed by RL, then evaluated on 100 articles with publicly available XML ground truth that none of the models saw in training.
Benchmark Results
We see that both RL and SFT manage to improve the benchmark performance, with RL after SFT leading to continued improvement.
Some of the score increase comes from better fluency with the tools and requirements of the JATS agent profile. However, while digging through the traces another reason became clear: both the base and SFT models get stuck in tool calling loops that block them from finishing the task.
Countering Looping
A typical SFT loop: the model patches a script, gets an error, and sends the exact same command again, up to 41 times in a row. Often its own reasoning flags the problem (“I should consider using sed”) right before it repeats the call. RL gets rid of almost all of this, whichever model it starts from.
Greedy Decoding
Inference at temperature = 0 is a common way to make an agent behave more predictably, so we reran the benchmark that way. It made results significantly worse for non-RL models.
Greedy decoding hurts every model except RL-only. SFT finishes just 9% of articles, repeating the same malformed tool call each turn. Some of this behavior arises at temperature 0.7, yet sampling usually breaks it out after one bad call. SFT-then-RL improves but still tends to fall back into identical-call loops.
Root Causes
The SFT model didn’t learn to loop from its teacher: none of the teacher traces repeats a call more than three times in a row.
Our strongest hypothesis is that the main driver is the off-policy nature of SFT: the model learns only from the teacher’s trajectories, never its own. The teacher rarely breaks its own scripts, so when the student does, it’s in a state the training data never showed. With no example of “failed edit, then a different approach”, the highest-probability continuation is whatever is already in context: the previous command. This is the textbook compounding-error, or exposure-bias, failure of imitation learning.
RL fixes this because it is on-policy: the model trains on its own rollouts, so it can adapt to its own mistakes. A rollout that loops to the turn cap delivers nothing and scores 0, while one that hits the same error and tries something else can still get a positive advantage. That contrast is exactly what the update learns from. Over 40 steps, rollouts with 4 or more identical calls have a negative advantage 88% of the time and drop from 12% to 4%.
To isolate the on-policy part, we also tried on-policy distillation (OPD). Like SFT, it learns from a teacher, but it trains on the student’s own runs: the student generates the trajectory, and the teacher grades each token it produces.
OPD matches RL’s benchmark score at temperature 0.7 and loops less than SFT. At temperature 0, it holds up where SFT collapses. It still loops occasionally at temperature 0, so OPD isn’t quite as effective as RL at eliminating loops entirely. Still, the evidence points clearly to the same conclusion: a student that learns from its own trajectories performs better and loops less.
What’s next
This was a small experiment, but it’s shaped how we think about training models for our agents and the importance of on-policy data. We’re excited to keep training models for more document tasks and will keep sharing what we learn.
If you’re training agents and have run into similar (or different!) issues, we’d love to compare notes.