Latest
Subscribe

Why Keeping an AI's Knowledge Current Means Choosing Between Retrieval and Retraining

The same underlying question, updating what a model knows, produces two very different engineering answers, and the difference matters most when facts change.

a close up of a network with wires connected to it
Photo · Photo by Albert Stoynov on Unsplash

The problem underneath the jargon

A large language model learns its knowledge during training, then that knowledge is frozen. Ask it about something that happened after its training cut-off, or about a private document it never saw, and it will either say it does not know or, worse, guess confidently and get it wrong. Two different techniques exist to fix this: retrieval-augmented generation (RAG) and fine-tuning. Both are ways of giving a model information it did not originally have. But they solve the problem in fundamentally different places, and that difference has practical consequences for anyone relying on the output.

What retrieval actually does

RAG does not change the model at all. Instead, when a question comes in, a separate search system first fetches relevant text, a policy document, a product manual, a set of recent news articles, from a database, and hands that text to the model alongside the question. The model then answers using what it was just shown, in addition to what it already knew from training. Think of it as an open-book exam: the model’s underlying knowledge and reasoning ability are unchanged, but it now has a specific, current set of notes in front of it.

The practical upshot is that updating the system means updating the database, not the model. Add a new document, and the next query can immediately draw on it. There is no retraining, no waiting, no computing cost beyond running the search. This is why RAG is the default choice for anything that changes often: pricing pages, internal policies, regulatory guidance, live news.

What fine-tuning actually does

Fine-tuning is different in kind, not degree. It takes an already-trained model and continues training it on a smaller, targeted set of examples, adjusting the internal parameters, the numerical weights, that determine how the model responds. The knowledge or behaviour becomes baked into the model itself. Nothing needs to be fetched at question time, because the model has, in effect, absorbed the material.

This is closed-book learning. It suits situations where you want to change how a model behaves, its tone, its format, its way of following instructions, or its fluency in a specialised style, rather than simply what facts it can access. It is also the right tool when you need the model to reliably follow a particular structure or reasoning pattern across many similar tasks, something a document stuffed into a prompt cannot easily teach.

Where the two genuinely diverge: keeping things current

The clearest practical difference shows up over time, as facts change. With RAG, updating knowledge is close to instant and reversible: swap out a document, and the system’s answers change on the next query. Nothing is unlearned because nothing was learned in the model itself. With fine-tuning, updating knowledge means running a new training pass, which costs computing time and money, and there is a real risk of what researchers call catastrophic forgetting, where teaching the model new material degrades its performance on things it previously handled well. A fine-tuned model showing its age cannot simply have a file swapped out; it typically needs to be retrained.

This matters enormously for anything where facts have a shelf life: interest rates, legal thresholds, product specifications, safety guidance. An organisation relying on a fine-tuned model for this kind of information is committing to an ongoing retraining cycle. An organisation using RAG is committing to keeping its document store accurate, which is usually cheaper and faster.

Where the two genuinely diverge: trust and traceability

The second major difference is verifiability. Because RAG shows the model a specific passage before it answers, well-built systems can cite exactly which document a claim came from, letting a human check the source. This is valuable wherever accountability matters, in professional, regulatory or safety-critical settings. Fine-tuning offers no equivalent. Once knowledge is absorbed into the model’s weights, there is no way to point to “this is where that fact came from”; the model simply produces an answer shaped by everything it was trained on, with no built-in citation trail.

They are not mutually exclusive

In practice, many production systems use both. A model might be fine-tuned to follow a particular format, tone or task structure, and then given retrieval on top of that so its factual content stays current and checkable. The distinction to hold onto is what each technique changes: fine-tuning changes the model itself, permanently and at a cost each time it is repeated; retrieval changes what the model sees at the moment it answers, cheaply and repeatedly. Anyone evaluating an AI tool that claims to “know” something recent or specific is worth asking a simple question: is that knowledge sitting in a document the system can point to, or is it baked into the model with no paper trail at all.

For readers wanting to understand the underlying model-training process itself, or the broader question of which approach suits which task, further detail sits with the standards bodies and research organisations tracking this fast-moving field.

Sources