Hindsight Structured Advice Distillation: Teaching a Small Model to Guide a Research Agent
A research agent's answer depends on more than the model behind it. The same agent, with the same tools and the same data, can write an excellent answer to one question and a thin answer to the next. Often the difference is know-how about the task: which filings carry a given figure, that some companies report product lines only as a share of sales, or that a quarterly series has to be rebuilt from year-to-date filings.
There are two usual ways to supply this know-how. Retraining the agent is slow and expensive, and every change to a production system has to be validated again. Writing the know-how into the agent's prompt, perhaps optimized automatically [1, 2, 3], works up to a point, but the prompt grows long and is the same for every question. An instruction that helps a data-extraction question, such as "present results in a table", can make a qualitative question worse.
We take a third route and learn to generate the guidance itself, building on recent work that trains a small model to write guidance for a larger one [4, 5]. The advisor reads each question and writes a plan that fits it. The plan is added to the agent's instructions, and the agent then researches and answers exactly as it normally would. The agent is never retrained. The advisor is a small model, so it is cheap to train and adds little latency.
In this post, we describe how we train this advisor from the agent's own history of successes and failures, a method we call hindsight structured advice distillation. On FrontierFinance [6], our public benchmark of expert-written research questions, the advised agent scores 56.6%, up from 52.9%. That puts it above the agent's own high-effort configuration and above every frontier model on the leaderboard, at 37% lower cost than the high-effort configuration. The same recipe also carries over to other domains: with a small open agent, it improves online shopping, text-to-SQL and expert document extraction on three public benchmarks.
The idea builds on a line of research called distillation with privileged information [7, 8]. A teacher model is given information that will not be available later, such as the correct answer, and produces better outputs with it. A student model then learns to produce those outputs without the extra information. Recent work uses this recipe to train the agent itself [9, 10]. We use it to train the advisor instead.
In our case the privileged information is hindsight. After the fact, on a question with a graded answer, it is easy to see what would have helped. An answer that gave annual figures when the question asked for quarters shows that the agent needed to be told to cover every period. We turn this kind of hindsight into training data in four steps.
1. Study the failures. We run the agent on 1,819 training questions and grade each answer against its rubric. These questions are a separate, proprietary set, collected with the same expert-annotation process as FrontierFinance, and none of them is among the 220 public queries we evaluate on. For every weak answer, a strong model writes down, in plain language, why it fell short. We then group these notes into a checklist of 103 recurring gaps and fixes, such as "enumerate all periods in the requested range" or "pull sub-segment detail within reported segments".
2. Try different strategies. Separately from the checklist, we write three whole-answer strategies, such as "prefer structured data sources" or "extract figures from filing tables", and also keep the agent's default behaviour. We answer every training question these four ways and record which one scored best.
3. Write the ideal plan. A teacher model writes the plan that would have helped most. It sees the question and its rubric, all four answers with their scores and tool usage, and the full tool-call trace of the winning strategy. It also gets the menus to choose from: the four strategies, the data sources, and the 103-item checklist from step 1. Because the teacher only picks from these fixed menus, the plan cannot give away the expected answer.
4. Train the advisor. We train a small open model, Qwen3-8B [11], with supervised fine-tuning (SFT) to write that plan from the question alone, since the question is all it will see in production. Each training example pairs a question with the teacher's plan for it.
The plan is a short structured form rather than an essay. Every field takes its value from a fixed menu, which keeps the advice predictable, easy to inspect and cheap to generate. Newer decision-style models such as Jev [12], which pick among declared options in a single fast step, could make generating it faster still.
FrontierFinance [6] is our open benchmark of 220 expert-crafted research queries, graded against 11,543 source-attributed rubric items across six investor use cases. The public leaderboard compares systems on the share of rubric items their answers satisfy, along with cost and latency per query. We ran the advised agent on the same 220 queries with the same grader.
The advice makes the agent do somewhat more work, which raises cost by about 23% over the unadvised agent. That is still well under the high-effort configuration, which uses better tools, stronger context management and more reasoning, and which the advised agent slightly outscores. We see a similar gain on a separate internal validation set, so the result is not specific to the public queries.
To see which part of the plan does the work, we removed one field at a time from the teacher's plan and measured the change in quality. The watch-outs from the checklist carry nearly all of it. Without them, the gain disappears entirely. Sources, the tool plan, web use and length each add a smaller amount. This fits how the checklist was built. Each item came from studying a real failure of this agent, so the advice points the agent at gaps it actually has.
The two examples below come from the public test set. In each, the advisor saw only the question.
In the Colgate-Palmolive question, two of the watch-outs describe the situation the agent ran into. One asks it to tell apart a figure that is missing from one that is zero, and the other a disclosure that is stated from one that is only implied. In the ASML question, the plan chose a table-first strategy, allowed web search as a fallback, and pointed the agent to transcripts and news as well as filings. Its watch-outs asked for regional detail and capital returns alongside guidance, which is what an investor expects from an earnings summary.
To see whether the recipe carries over, we applied it with a small open model, Qwen3-4B [12], as the agent on three public benchmarks. In online shopping (WebShop [13]), the score rose from 54.6 to 63.4 and the share of exactly correct purchases from 3.8% to 23.2%. There the advisor's job turned out to be simple. It copies the options the shopper asked for, such as color and size, into the plan as exact buttons to click, which the small agent kept missing on its own. On database questions (BIRD [14], text-to-SQL), accuracy rose from 40.0% to 43.1%. On expert document extraction (ExpertLongBench [15]), coverage of the expert checklist rose from 49.7% to 54.5%.
The advice itself is flexible. It is added to the agent's instructions as a suggestion, and the agent is free to deviate from it when the evidence points elsewhere. Some fields describe the question rather than the agent, such as which sources hold the answer, how many periods it spans, or which disclosures to check, and those carry over naturally to other agents. Other fields reflect the habits of the guided agent, such as its typical blind spots, and are most useful for that agent. Across many questions, the advice helps on average even though the agent does not follow every suggestion.
A production agent can be improved with guidance generated by a small advisor model, without retraining the agent. On FrontierFinance, the advisor puts our agent at the top of the public leaderboard and reaches the quality of its high-effort configuration at 37% lower cost per query, and the same recipe carries over to three other domains. It is a general way to improve a frozen agent, and we are in an early testing phase of this approach, with a plan to roll it out to production. We draw three conclusions from this work.
1. Hindsight is a rich source of training data. Running an agent on past questions, seeing what worked, and writing down what would have helped gives a small model something concrete to learn.
2. The best advice comes from the agent's own failures. The watch-outs that carry most of the gain were mined from this agent's weak answers. Many of them describe what a good answer to the question needs, so they are useful well beyond the one agent they were learned from.
3. A small guidance model is a cheap adapter. It supplies task know-how that a general agent lacks, question by question, without retraining the agent or growing its prompt. A short form filled from fixed menus is easy to audit and inexpensive to produce, and decision-style models could make it cheaper still.
We thank Yuhao Zhang and Ashwin Paranjape for their comments on this work.
1. Yongchao Zhou et al. Large Language Models Are Human-Level Prompt Engineers. International Conference on Learning Representations (ICLR), 2023. arXiv:2211.01910.
2. Omar Khattab et al. DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines. International Conference on Learning Representations (ICLR), 2024. arXiv:2310.03714.
3. Lakshya A Agrawal et al. GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning. International Conference on Learning Representations (ICLR), 2026. arXiv:2507.19457.
4. Zekun Li, Baolin Peng, Pengcheng He, Michel Galley, Jianfeng Gao and Xifeng Yan. Guiding Large Language Models via Directional Stimulus Prompting. Advances in Neural Information Processing Systems (NeurIPS), 2023. arXiv:2302.11520.
5. Parth Asawa et al. How to Train Your Advisor: Steering Black-Box LLMs with Advisor Models. International Conference on Machine Learning (ICML), 2026. arXiv:2510.02453.
6. Yuhao Zhang, O. Ozan Koyluoglu, Thejas Venkatesh, Richard Diehl Martinez, Vishank Bhatia, Arash Alidoust and Ashwin Paranjape. FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents. arXiv:2608.11683, 2026.
7. Vladimir Vapnik and Akshay Vashist. A New Learning Paradigm: Learning Using Privileged Information. Neural Networks, 22(5–6):544–557, 2009.
8. David Lopez-Paz, Léon Bottou, Bernhard Schölkopf and Vladimir Vapnik. Unifying Distillation and Privileged Information. International Conference on Learning Representations (ICLR), 2016. arXiv:1511.03643.
9. Dian Chen, Brady Zhou, Vladlen Koltun and Philipp Krähenbühl. Learning by Cheating. Conference on Robot Learning (CoRL), 2019. arXiv:1912.12294.
10. Emiliano Penaloza, Dheeraj Vattikonda, Nicolas Gontier, Alexandre Lacoste, Laurent Charlin and Massimo Caccia. Privileged Information Distillation for Language Models. International Conference on Machine Learning (ICML), 2026. arXiv:2602.04942.
11. An Yang et al. Qwen3 Technical Report. arXiv:2505.09388, 2025.
12. Diogo Almeida. Introducing System One Models & Jev. TypeSafe AI blog, 2026.
13. Shunyu Yao, Howard Chen, John Yang and Karthik Narasimhan. WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents. Advances in Neural Information Processing Systems (NeurIPS), 2022. arXiv:2207.01206.
14. Jinyang Li et al. Can LLM Already Serve as A Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQLs. Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track, 2023. arXiv:2305.03111.
15. Jie Ruan et al. ExpertLongBench: Benchmarking Language Models on Expert-Level Long-Form Generation Tasks with Structured Checklists. International Conference on Learning Representations (ICLR), 2026. arXiv:2506.01241.