ResearchRAG
How Fin learns which Guidance, Data Connectors, and Procedures matter for each conversation

How Fin learns which Guidance, Data Connectors, and Procedures matter for each conversation
Customers can configure Fin to handle conversations the way their support team would: applying policies, checking systems, and following processes specific to their business.
They can do this themselves, in three ways:
Take an online retailer:
As Fin moves deeper into a support operation, the configuration around it naturally grows. What starts as a few instructions for common cases becomes a wider map of policies, systems, exceptions, and processes. Over time, the configuration comes to reflect how the business actually runs.

For the retailer, a “where’s my order?” conversation needs a narrow slice of its configuration items: the order-lookup connector and the guidance that applies to delivery. Gift cards, subscriptions, and the damaged-item Procedure can stay out of view.
The simplest implementation is to put every configured behaviour in front of the agent on every turn and let it determine what applies. This works while the configuration is small. But as the business adds more policies, systems, and processes, the prompt keeps growing even though each conversation still needs only a small subset of them.
The agent is forced to process excess input, filter through irrelevant configuration, and consider actions completely unrelated to the request. This context noise makes target behaviour harder to isolate, all while inflating compute costs and adding delay to every turn. Switching to a larger, more expensive model to rescue instruction-following only makes that cost and latency footprint worse.
At scale, an agent’s capability depends on selecting the right context for each conversation. Provide the Guidance, Data Connectors, and Procedures that matter, and leave everything else out.
The system prompt has become a retrieval problem.

We solve this with runtime retrieval: for each conversation, Fin selects the relevant Guidance, Data Connectors, and Procedures and uses them to assemble the prompt.

This extends the logic of RAG from knowledge to behaviour. RAG retrieves the content a model needs to answer; Fin retrieves the configuration items that determines how the agent answers, which systems it can use, and which conversation path is followed.
That distinction raises the cost of a miss. A missed rule can alter the response, a missed connector can prevent an action, and a missed procedure can send the conversation down the wrong path.
Hence, the greatest challenge is learning relevance across customer configurations that are unique to each workspace and change continuously.
Fin already produces those relevance judgments in the course of serving customers. When the answer model writes a reply, it cites the Guidance it used. When the planner decides what to do, it scores the relevance of each Data Connector and Procedure. Each conversation therefore leaves behind a record of which parts of the configuration mattered.
We distill those decisions into training labels for smaller relevance models that run earlier in the serving pipeline.
For Guidance, a cited conversation–rule pair becomes a positive example, while Guidance that was available to the answer model but went unused becomes a negative. Data Connector and Procedure labels come from the planner’s relevance decisions.
Because we build Fin’s models and serving stack end to end, the production models that serve customers can also act as teachers: their judgments become training data at scale, under strict eligibility and consent controls.
The Guidance selector is trained on roughly three million conversation–guideline pairs distilled from production traffic. The Data Connector and Procedure selectors are trained on more than one million planner decisions drawn from a single week of conversations.
Each selector is a cross-encoder fine-tuned from Gemma 4 E2B. It processes the live conversation and one candidate instruction together, then predicts whether the production model found that instruction relevant.
The training objective is binary relevance. We apply low-rank adapters across the attention and MLP layers using LoRA with rank 16 and alpha 32, while training the scoring head in full. The adapters are then merged into the base model for serving.
The production models supervise the selectors, and the selectors in turn sharpen the context those models receive.
When fresh labels are needed, selection is disabled for a small fraction of traffic. The production models see the full configuration again, and their unfiltered decisions become the next round of training data.

Guidance, Data Connectors, and Procedures each get their own selector — the same cross-encoder recipe, trained separately on each surface’s labels. They also each need their own contract for what the selector is allowed to remove, because they fail differently: a guidance rule affects what Fin says, a connector affects what Fin does, and a procedure decides which path the conversation takes.

Guidance: shortlisting. For any conversation, the selector keeps the guidance most relevant to what the customer is asking and leaves the rest out. It activates only once a workspace’s guidance grows past a set threshold, and it fails open: on any error or timeout, the full set passes through unchanged, so retrieval never becomes a new source of risk.
Data Connectors: deciding when planning is worth it. Connectors are where Fin does work, and deciding what to do is one of the most expensive steps in the pipeline. Yet most conversations are straightforward questions that need no connector at all — in production traffic, 61% of planning calls came back with an empty plan. Connectors also chain: one call often feeds the next, so removing candidates one by one could break a plan that needed them together. The safe contract is a gate rather than a shortlist. If anything looks relevant, the planner runs with the full set; if nothing does, the planner is skipped entirely. The gate is tuned to protect recall first, because missing a real request costs far more to customer operations than an unnecessary planning call.
Procedures: letting “none” be an answer. Procedures don’t chain the way connectors do — Fin picks one and follows it — so they can be shortlisted safely. What makes them different is how often the right answer is no procedure at all. Most conversations don’t need one: in production traffic, 84% of procedure relevance checks find nothing that applies. That is why the selector uses an absolute relevance bar rather than a fixed top-N: a top-N list always returns something, even when nothing applies, while a threshold lets Fin conclude that this conversation doesn’t need a procedure and move on.


We evaluated each selector in two steps: first offline on held-out production traffic, and then in live A/B tests with the full-prompt behaviour as the control.
To measure selector performance before deploying to live traffic, we evaluate offline on a held-out dataset spanning 1 week of production traffic across all surfaces.
In this setup, decisions made by the full-prompt production model serve as the gold ground-truth labels. The benchmark evaluates how faithfully each selector preserves the context chosen when the production model sees the complete configuration. Performance is measured using two key metrics:
Because each surface tolerates retrieval failure differently, the selection mechanisms and operating points vary:
| Metric | Guidance selector | Data Connector selector | Procedure selector |
|---|---|---|---|
| Operating point | Top-10 shortlist | Calibrated threshold | Calibrated threshold |
| Recall | 0.97 | 0.95 | 0.95 |
| Filter rate | 0.57 | 0.44 | 0.74 |
Recall and filter rates describe the selectors, not the support experience. The measure that matters in production is whether Fin still resolves the conversation, so the A/B tests tracked resolution rate — the share of conversations Fin resolves without human intervention — alongside answer quality, cost, and latency.
Resolution rate remained broadly stable across the three tests. For the Data Connector and Procedure selectors, it was unchanged. The Guidance selector produced a small decrease, but also revealed the most interesting result: Fin adhered more closely to the behaviour customers had configured.
Citation rate, the share of replies that applied a configured rule increased because relevant Guidance was easier for Fin to surface and act upon. However, following Guidance faithfully often required Fin to ask clarifying questions or carry out additional steps, rather than resolving the conversation in a single turn.
This exposes an interesting distinction: maximising resolution rate is not necessarily the same as maximising adherence to customer-configured behaviour. A small reduction in immediate resolution can reflect Fin following instructions more faithfully, rather than a decline in answer quality.
The Guidance selector also reduced cost per conversation while answer quality held steady. For Data Connectors and Procedures, both cost and latency fell: when no planning was needed, a step that previously took seconds was replaced by a scoring pass that took milliseconds.
We are not sharing exact effect sizes for competitive reasons. Taken together, the selectors allowed Fin to read substantially less configuration, follow it more precisely, and operate at lower cost and latency with resolution unchanged for Data Connectors and Procedures, and a small, explainable trade-off for Guidance.
For Fin, prompt engineering now includes deciding what enters the prompt on every turn. We made that decision a learned retrieval step: trained from production, evaluated end to end, and designed around how each type of behaviour works.
Fin can therefore scale with the policies, systems, and processes of each customer’s business while maintaining answer quality, following configured behaviour more faithfully, and reducing cost and latency.