BlogBuilding Fin
Hardening Fin’s billing systems with a self-maintaining knowledge base.

Over the last 16 months, R&D at Fin has used AI to 3x productivity. While teams working on the core product have continued to accelerate and push the boundaries of AI-enabled development, the monetization teams that handle pricing, subscriptions, and invoicing (a domain collectively known as “quote-to-cash”), have had less room to experiment. In our corner of the product, bugs incur significant real-world cost: inaccurate invoices, customers charged for something they didn't buy, or people locked out of a product they're entitled to. Work in this area is high-stakes; it relies on billing expertise and historical context, and every change carries some degree of risk.
With these constraints in mind, we began working on systems that would also allow us to increase productivity with AI. The goal was to let anyone, human or AI Agent, change billing code with confidence and to use AI to build guardrails for:
Passing the right domain context to the AI was the prerequisite, and it came with two challenges. First, the critical knowledge was scattered and unstructured. Second, more context does not mean better answers: pass too much of it and the LLM's reasoning degrades in the noise. Before any guardrail could work, we had to solve both.
Billing is context-heavy. It's guarded by rules and regulations, entangled with billing providers, CRM platform, quoting tools, and details that are only obvious once you've spent enough time in the domain. The critical knowledge (how subscriptions interact with upstream systems, what a webhook is allowed to do, which customer entitlement states are correct) was spread across the codebase, docs, GitHub threads, runbooks, past incidents, and domain experts' heads.
An AI Agent reviewing a billing pull request without that context can only reason based on general knowledge: how billing systems typically work, the common Rails patterns, sound engineering practice and whatever it gleans from the codebase.
Pointing the model at everything we had was inadequate: dumping every doc, thread, and runbook into the context window dilutes the attention of the AI. What we needed was a way to give the Agent the minimum relevant context for the task at hand, a form of retrieval-augmented generation.
We started by putting AI to work ingesting all the scattered sources and distilling the information into a knowledge base of structured, retrieval-ready documents – a variation of how Meta mapped tribal knowledge in its data pipelines:
The full source documents (architecture docs, flow descriptions, runbooks) live alongside them in the knowledge base, as reference material the distilled rules link back to.
Here's an example – a rule that only exists because of a quirk in the billing provider's API:

The distillation itself is packaged as reusable AI skills. Hand one a source document and it turns it into distilled, retrieval-ready entries, so growing the knowledge base is a low-effort operation.
We kept the retrieval deliberately simple with no embeddings or vector database. The Agent searches the knowledge base the way an engineer would: scan an index, grep for terms, and read the matching files. That also keeps the retrieval transparent – when AI pulls the wrong context, you can see exactly why. And because each document is small and self-contained, five to 15 of them are enough to ground AI without flooding its context window. The front door to all of it is a single auto-generated index file, a wiki-style map with one line per document containing its ID and a one-sentence summary. Rules are grouped by the billing entity they describe (subscription, invoice, contract, amendment, customer) and past investigation patterns by root cause (config errors, code bugs, timing issues).
The search terms come from the task itself: entity names, billing object types, or interaction keywords. A PR review and an incident investigation each pull their own thin slice of the knowledge base.
At its current size of a few hundred documents, an index the model can read whole is the most efficient solution. The model picks context by reasoning, not just similarity. Once the size of the knowledge base proves challenging or our evaluations start showing that relevant documents aren't being retrieved, semantic search may be worth building.
We seeded the knowledge base with a one-off AI pass over our existing sources. From there, automated AI jobs keep it current and keep score:
This is how we teach AI the deep domain without retraining the model. The learnings from every on-call ticket and every incident flow back into the knowledge base, so every case the system gets wrong makes it better at the next one. This keeps the manual toil of managing this system to a minimum, but a human stays in the loop. Every edit to the knowledge base arrives as a pull request for review.

Curated domain context is what makes AI-assisted workflows across different stages of development truly powerful.
Every billing pull request is automatically reviewed against the relevant knowledge base content. The AI reviewer cites documented rules, leaving comments like:

This supplements the engineer's diligence. Nobody holds every provider constraint and edge case in their head; a reviewer with the whole knowledge base at its disposal catches what you didn't know to check.
Every AI-generated comment is verified and graded:
This is the piece we're most excited about. Domain-aware review guards the changes being made today, but plenty of production state wasn’t shaped by today's code. Our data carries years of history like migrations from earlier billing systems, data contracts with CRM and billing providers that can evolve, manual concessions applied for individual customers, and the occasional race between concurrent changes. A change can pass review and every test, and still meet a customer record produced by a code path that was retired three years ago. So this guardrail doesn't watch code paths at all; it audits the accumulated state itself, every day.
It does this by monitoring invariants, statements about a customer state that must always be true. For example: "every subscription bills for everything the customer is entitled to." We can write invariants in plain English because the AI has the domain context (what a subscription is, what its subtypes are, when an exception legitimately applies), and can compile that sentence into a precisely scoped daily audit query against our data warehouse.
This invariant:
A subscription must never contain duplicate items: at most one active recurring item for each combination of product and role.
compiles into the core predicate of a generated SQL audit:

When a violation of an invariant is detected, it alerts the engineering Slack channel, naming the exact customer.
Initially, every invariant was written by an engineer, usually after a deliberate review of the code or an incident. We now use AI Agents to discover these rules automatically. An Agent reads the full subscription history of a sample of customers from our data warehouse, together with the knowledge base and the code paths that produced those records. From that it works out the states, transitions, and patterns that govern the system, and proposes them as candidate invariants for an engineer to review.
Every on-call ticket now arrives enriched not only with relevant logs, traces, and the customer's subscription history, but also with the domain context. The Agent also sanity-checks the ticket itself. It verifies the issue is routed to the team that owns the affected code, suggests an urgency and a status, and determines if it’s a genuine on-call issue, a feature request, or project work in disguise. The root-cause hypothesis and suggested next steps carry a confidence level based on how much evidence supports them. A real example, condensed:

After the on-call ticket is closed, the AI-suggested root cause is graded based on:
Within six weeks of launch, domain-aware code review had flagged more than 100 issues on pull requests, each fixed before the change merged. A growing share of code review comments also cite a specific documented rule, signalling the knowledge base is working.
We've defined over 100 invariants so far, with the daily audits built from them already catching inconsistencies in long-standing customer records.The tooling is also creating new habits: the question that now ends most incident reviews is "should this be an invariant?"
Behind each of those catches is a scenario we set out to prevent: an inaccurate invoice reaching a customer, an account charged for something it didn't buy, or a workspace locked out of a product it's entitled to.
Auto-triage is the most challenging guardrail to get right: the evaluation data is still thin, and on-call ticket inflow is noisy enough that we can't yet draw a clean before/after trend line. What we can see is the direction: resolution times are trending down, and we've received the first indications from the engineering teams that the AI-generated root-cause hypothesis proves helpful in investigations.
Overall, we’re getting faster. Pull requests get domain-expert review in minutes and on-call tickets arrive pre-investigated. The next gear is removing steps entirely. We're not auto-resolving incidents on money-bearing code paths yet, or auto-approving PRs the way other engineering teams at Fin already do, but every graded triage and every useful domain-grounded comment on a pull request is evidence accumulating toward that. When we do hand the AI more autonomy, it'll be a decision backed by a track record we can measure.