What is a RAG system?
A RAG system (retrieval-augmented generation) is an AI system that searches a dedicated knowledge base for relevant passages before answering, and formulates its answer solely on that basis. Instead of answering from training knowledge, the system looks things up first – like a person consulting the manual. That makes answers current and verifiable: every statement can be traced back to a specific document.
How a RAG system works
1. Preparation: your documents are split into meaningful passages and converted into numerical vectors (embeddings) that represent their meaning. Those vectors go into a vector database.
2. Retrieval: when a question arrives, it too is converted into a vector. The system looks for the passages closest in meaning – semantically, not by keyword. A question about “holiday entitlement” will therefore also find a passage about “annual leave”.
3. Generation: the passages found are handed to the language model together with the question, with the instruction to answer solely from them. The result is a formulated answer including references to the passages used.
The decisive difference from a plain chatbot: the model does not have to “know” anything. New documents are available immediately, without the model being retrained – and outdated content disappears as soon as it is removed.
Why RAG suits company knowledge
Every answer points to its source – verifiable rather than a black box.
New documents are indexed and usable straight away, with no retraining.
The model answers from the passages presented to it rather than from vague training knowledge.
Internal manuals, policies and contracts become searchable – content no public model knows.
The content stays in your knowledge base; it does not have to be trained into a model.
Combined with a permissions concept, only sources the person asking may see are drawn on.
A RAG system can only be as good as its sources: if the underlying documents are wrong, outdated or contradictory, the system reproduces those faults accordingly. It does not check content for factual accuracy. That is why the reference matters – and why checking answers professionally remains the user's job. RAG does not make AI infallible, but it does make it traceable.
Frequently asked questions
What is the difference between RAG and fine-tuning?
Fine-tuning trains a model further on additional data – costly, and new content requires training again. RAG leaves the model unchanged and supplies knowledge at runtime. For company knowledge that changes frequently, RAG is therefore usually the better fit.
Does RAG need a vector database?
In practice almost always: it enables semantic search, that is finding by meaning rather than by exact keyword. Some systems additionally combine it with classic full-text search to improve precision and recall.
Does RAG prevent hallucinations completely?
No, but it reduces them considerably, because the model answers from concrete passages rather than from memory. A residual risk remains – for instance when sources contradict each other. The reference does make such cases checkable, though.
How do our documents get into the system?
They are uploaded and prepared and indexed automatically. With KOSMO, new content is available immediately afterwards – with no retraining and without the knowledge base having to be rebuilt.
What this looks like with KOSMO
Theory is one thing – in 30 minutes we show you live how KOSMO does this in your organisation. With your own content.







