← Back to Blog

Hindsight Structured Advice Distillation: Teaching a Small Model to Guide a Research Agent

Ozan Koyluoglu, PhD/MBA

A research agent's answer depends on more than the model behind it. The same agent, with the same tools and the same data, can write an excellent answer to one question and a thin answer to the next. Often the difference is know-how about the task: which filings carry a given figure, that some companies report product lines only as a share of sales, or that a quarterly series has to be rebuilt from year-to-date filings.

There are two usual ways to supply this know-how. Retraining the agent is slow and expensive, and every change to a production system has to be validated again. Writing the know-how into the agent's prompt, perhaps optimized automatically [1, 2, 3], works up to a point, but the prompt grows long and is the same for every question. An instruction that helps a data-extraction question, such as "present results in a table", can make a qualitative question worse.

We take a third route and learn to generate the guidance itself, building on recent work that trains a small model to write guidance for a larger one [4, 5]. The advisor reads each question and writes a plan that fits it. The plan is added to the agent's instructions, and the agent then researches and answers exactly as it normally would. The agent is never retrained. The advisor is a small model, so it is cheap to train and adds little latency.

table1
QuestionWhat is Colgate-Palmolive's revenue by geography and product category for each quarter since 2019?
AdviceCover every quarter in the range. If a figure is not disclosed, say so instead of estimating it. Separate what the company states from what can only be inferred.
Answer without adviceFills the product-category columns with estimates computed from percentage shares.
Answer with adviceReports quarterly revenue by region and states that product categories are disclosed only as a share of sales.

Table 1: The agent's answer without and with advice.


In this post, we describe how we train this advisor from the agent's own history of successes and failures, a method we call hindsight structured advice distillation. On FrontierFinance [6], our public benchmark of expert-written research questions, the advised agent scores 56.6%, up from 52.9%. That puts it above the agent's own high-effort configuration and above every frontier model on the leaderboard, at 37% lower cost than the high-effort configuration. The same recipe also carries over to other domains: with a small open agent, it improves online shopping, text-to-SQL and expert document extraction on three public benchmarks.

fig1_leaderboard
FrontierFinance: quality against cost and latency

220 public queries. Hover a point for its name.

Samaya System with advisorSamaya harnessFinance Agent v2 harnessWeb-search harness

Figure 1: Rubric qualification rate against cost and latency per query. All points except the advised agent are from the public FrontierFinance leaderboard; the advised agent's run used the benchmark's queries and grader.

Learning from hindsight

The idea builds on a line of research called distillation with privileged information [7, 8]. A teacher model is given information that will not be available later, such as the correct answer, and produces better outputs with it. A student model then learns to produce those outputs without the extra information. Recent work uses this recipe to train the agent itself [9, 10]. We use it to train the advisor instead.
‍
In our case the privileged information is hindsight. After the fact, on a question with a graded answer, it is easy to see what would have helped. An answer that gave annual figures when the question asked for quarters shows that the agent needed to be told to cover every period. We turn this kind of hindsight into training data in four steps.

fig2_method
How the advisor learns, and how it is used

Training happens once, offline. At test time the agent is unchanged; it only receives a short plan.

01Hindsight
Study failures
Run the agent on past questions, grade every answer, and note why weak answers fell short. The notes become a checklist of recurring gaps.
→
02Hindsight
Try strategies
Answer every question several ways and record which way actually scored best.
→
03Structured advice
Write the ideal plan
A teacher that can see the grading and every answer writes the plan that would have helped, as a structured form.
→
04Distillation
Train the advisor
A small model learns to write that plan from the question alone.
At test time
Question
→
Advisor writes a plansmall model, one short call
→
Samaya agent answersquestion + plan, same tools

Figure 2: Hindsight structured advice distillation. The advisor learns from the agent's graded history which plan would have helped, and at test time it writes that plan from the question alone.


1. Study the failures.
We run the agent on 1,819 training questions and grade each answer against its rubric. These questions are a separate, proprietary set, collected with the same expert-annotation process as FrontierFinance, and none of them is among the 220 public queries we evaluate on. For every weak answer, a strong model writes down, in plain language, why it fell short. We then group these notes into a checklist of 103 recurring gaps and fixes, such as "enumerate all periods in the requested range" or "pull sub-segment detail within reported segments".
2. Try different strategies. Separately from the checklist, we write three whole-answer strategies, such as "prefer structured data sources" or "extract figures from filing tables", and also keep the agent's default behaviour. We answer every training question these four ways and record which one scored best.
3. Write the ideal plan. A teacher model writes the plan that would have helped most. It sees the question and its rubric, all four answers with their scores and tool usage, and the full tool-call trace of the winning strategy. It also gets the menus to choose from: the four strategies, the data sources, and the 103-item checklist from step 1. Because the teacher only picks from these fixed menus, the plan cannot give away the expected answer.
4. Train the advisor. We train a small open model, Qwen3-8B [11], with supervised fine-tuning (SFT) to write that plan from the question alone, since the question is all it will see in production. Each training example pairs a question with the teacher's plan for it.
‍
The plan is a short structured form rather than an essay. Every field takes its value from a fixed menu, which keeps the advice predictable, easy to inspect and cheap to generate. Newer decision-style models such as Jev [12], which pick among declared options in a single fast step, could make generating it faster still.

table2
FieldWhat it tells the agentExample
StrategyWhich of four approaches to taketables_extraction
LengthHow long the answer should belong (3+ pages)
Web useWhether web search is neededavoid
SourcesWhich data sources to lean onFILINGS, …
Tool planWhich tools to call, roughly in orderlist_documents, search_documents
Watch-outsTwo to four items from the 103-item checklistenumerate all periods in the requested range

Table 2: The fields of the advisor's plan. Every field takes its value from a fixed menu.

Results on FrontierFinance

FrontierFinance [6] is our open benchmark of 220 expert-crafted research queries, graded against 11,543 source-attributed rubric items across six investor use cases. The public leaderboard compares systems on the share of rubric items their answers satisfy, along with cost and latency per query. We ran the advised agent on the same 220 queries with the same grader.

table3
SystemQualityCost / queryLatency / query
Samaya agent + advisor56.6%$1.14257 s
Samaya agent, high effort56.0%$1.81278 s
Claude Fable 5.1 (Finance Agent v2 harness)55.9%$6.89268 s
Samaya agent, no advice52.9%$0.93214 s
Claude Opus 4.8 (Finance Agent v2 harness)45.0%$2.61156 s

Table 3: FrontierFinance public leaderboard (220 queries).


The advice makes the agent do somewhat more work, which raises cost by about 23% over the unadvised agent. That is still well under the high-effort configuration, which uses better tools, stronger context management and more reasoning, and which the advised agent slightly outscores. We see a similar gain on a separate internal validation set, so the result is not specific to the public queries.

To see which part of the plan does the work, we removed one field at a time from the teacher's plan and measured the change in quality. The watch-outs from the checklist carry nearly all of it. Without them, the gain disappears entirely. Sources, the tool plan, web use and length each add a smaller amount. This fits how the checklist was built. Each item came from studying a real failure of this agent, so the advice points the agent at gaps it actually has.

fig3_ablation
What each part of the plan contributes

Drop in rubric items satisfied (points) when one field is removed from the teacher's plan. Internal validation set.

Watch-outs (checklist items)
6.7
Sources
4.4
Tool plan
3.0
Web use
2.7
Length
1.8

Figure 3: Leave-one-field-out ablation.

Two examples

The two examples below come from the public test set. In each, the advisor saw only the question.

fig4_examples
Two questions from the public test set

The advisor saw only the question. Scores are the share of rubric items satisfied.

“What is Colgate-Palmolive's revenue by geography and product category for each quarter since 2019?”
Advisor's plan
strategydefault
lengthlong (3+ pages)
web useavoid
sourcesFILINGS, …
Watch-outs
  • distinguish metric unavailability from zero
  • enumerate all periods in the requested range
  • derive standalone periods from cumulative filings
  • distinguish explicit from implicit disclosures
What changed

Without advice, the agent filled many quarterly cells with estimates, applying a percentage share to a total. With advice, it laid out the quarterly figures by region and said plainly that Colgate discloses product categories only as a share of sales, not in dollars.

36%
without advice
→
90%
with advice
“What is the summary of ASML's earnings?”
Advisor's plan
strategytables_extraction
lengthmedium (1–3 pages)
web usefallback
sourcesTRANSCRIPTS, FILINGS, NEWS
Watch-outs
  • search across all document types for the topic
  • surface management guidance and forward commentary
  • include geographic and segment breakdowns
  • include capital allocation and investment announcements
What changed

Without advice, the agent centred its summary on the latest quarter and its guidance. With advice, it also built a quarterly grid for the year and added the regional breakdown, including China, and the capital return programme.

47%
without advice
→
100%
with advice

Figure 4: The advisor's plan and its effect on two FrontierFinance queries. Field names are simplified and the tool plan is omitted.


In the Colgate-Palmolive question, two of the watch-outs describe the situation the agent ran into. One asks it to tell apart a figure that is missing from one that is zero, and the other a disclosure that is stated from one that is only implied. In the ASML question, the plan chose a table-first strategy, allowed web search as a fallback, and pointed the agent to transcripts and news as well as filings. Its watch-outs asked for regional detail and capital returns alongside guidance, which is what an investor expects from an earnings summary.

Beyond finance

To see whether the recipe carries over, we applied it with a small open model, Qwen3-4B [12], as the agent on three public benchmarks. In online shopping (WebShop [13]), the score rose from 54.6 to 63.4 and the share of exactly correct purchases from 3.8% to 23.2%. There the advisor's job turned out to be simple. It copies the options the shopper asked for, such as color and size, into the plan as exact buttons to click, which the small agent kept missing on its own. On database questions (BIRD [14], text-to-SQL), accuracy rose from 40.0% to 43.1%. On expert document extraction (ExpertLongBench [15]), coverage of the expert checklist rose from 49.7% to 54.5%.

fig5_domains
The same recipe in other domains

Qwen3-4B as the agent, held-out test sets. Grey is the agent alone; blue is the gain from advice.

Online shoppingWebShop, share of exactly correct purchases
3.8 → 23.2%
Database questionsBIRD text-to-SQL, execution accuracy
40.0 → 43.1%
Expert document extractionExpertLongBench, checklist coverage
49.7 → 54.5%

Figure 5: Results on three public benchmarks outside finance.


The advice itself is flexible. It is added to the agent's instructions as a suggestion, and the agent is free to deviate from it when the evidence points elsewhere. Some fields describe the question rather than the agent, such as which sources hold the answer, how many periods it spans, or which disclosures to check, and those carry over naturally to other agents. Other fields reflect the habits of the guided agent, such as its typical blind spots, and are most useful for that agent. Across many questions, the advice helps on average even though the agent does not follow every suggestion.

Conclusion

A production agent can be improved with guidance generated by a small advisor model, without retraining the agent. On FrontierFinance, the advisor puts our agent at the top of the public leaderboard and reaches the quality of its high-effort configuration at 37% lower cost per query, and the same recipe carries over to three other domains. It is a general way to improve a frozen agent, and we are in an early testing phase of this approach, with a plan to roll it out to production. We draw three conclusions from this work.

1. Hindsight is a rich source of training data. Running an agent on past questions, seeing what worked, and writing down what would have helped gives a small model something concrete to learn.
2. The best advice comes from the agent's own failures. The watch-outs that carry most of the gain were mined from this agent's weak answers. Many of them describe what a good answer to the question needs, so they are useful well beyond the one agent they were learned from.
3. A small guidance model is a cheap adapter. It supplies task know-how that a general agent lacks, question by question, without retraining the agent or growing its prompt. A short form filled from fixed menus is easy to audit and inexpensive to produce, and decision-style models could make it cheaper still.

Acknowledgments

We thank Yuhao Zhang and Ashwin Paranjape for their comments on this work.

References

1. Yongchao Zhou et al. Large Language Models Are Human-Level Prompt Engineers. International Conference on Learning Representations (ICLR), 2023. arXiv:2211.01910.
2. Omar Khattab et al. DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines. International Conference on Learning Representations (ICLR), 2024. arXiv:2310.03714.
3. Lakshya A Agrawal et al. GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning. International Conference on Learning Representations (ICLR), 2026. arXiv:2507.19457.
4. Zekun Li, Baolin Peng, Pengcheng He, Michel Galley, Jianfeng Gao and Xifeng Yan. Guiding Large Language Models via Directional Stimulus Prompting. Advances in Neural Information Processing Systems (NeurIPS), 2023. arXiv:2302.11520.
5. Parth Asawa et al. How to Train Your Advisor: Steering Black-Box LLMs with Advisor Models. International Conference on Machine Learning (ICML), 2026. arXiv:2510.02453.
6. Yuhao Zhang, O. Ozan Koyluoglu, Thejas Venkatesh, Richard Diehl Martinez, Vishank Bhatia, Arash Alidoust and Ashwin Paranjape. FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents. arXiv:2608.11683, 2026.
7. Vladimir Vapnik and Akshay Vashist. A New Learning Paradigm: Learning Using Privileged Information. Neural Networks, 22(5–6):544–557, 2009.
8. David Lopez-Paz, Léon Bottou, Bernhard Schölkopf and Vladimir Vapnik. Unifying Distillation and Privileged Information. International Conference on Learning Representations (ICLR), 2016. arXiv:1511.03643.
9. Dian Chen, Brady Zhou, Vladlen Koltun and Philipp Krähenbühl. Learning by Cheating. Conference on Robot Learning (CoRL), 2019. arXiv:1912.12294.
10. Emiliano Penaloza, Dheeraj Vattikonda, Nicolas Gontier, Alexandre Lacoste, Laurent Charlin and Massimo Caccia. Privileged Information Distillation for Language Models. International Conference on Machine Learning (ICML), 2026. arXiv:2602.04942.
11. An Yang et al. Qwen3 Technical Report. arXiv:2505.09388, 2025.
12. Diogo Almeida. Introducing System One Models & Jev. TypeSafe AI blog, 2026.
13. Shunyu Yao, Howard Chen, John Yang and Karthik Narasimhan. WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents. Advances in Neural Information Processing Systems (NeurIPS), 2022. arXiv:2207.01206.
14. Jinyang Li et al. Can LLM Already Serve as A Database Interface? A BIg Bench for Large-Scale Database Grounded Text-to-SQLs. Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track, 2023. arXiv:2305.03111.
15. Jie Ruan et al. ExpertLongBench: Benchmarking Language Models on Expert-Level Long-Form Generation Tasks with Structured Checklists. International Conference on Learning Representations (ICLR), 2026. arXiv:2506.01241.

‍

‍

‍