Understanding Its Purpose First

RAG first retrieves content relevant to the question from external sources, then passes it to the model to assist in generating a response.

Understanding Through an Example

When asked "What are the reimbursement requirements for our school's research activities?", the system first retrieves relevant passages from school policies, then organizes the answer based on them. Correct answers depend on whether relevant information was found, whether the data is up-to-date, and whether the model uses it correctly.

Visual explanationSee the Process Clearly
01Your question

What kind of evidence is needed

02Retrieve information

Retrieve relevant excerpts and sources

03Add context

Let the model reference specific content

04Generate and verify

Check if citations match conclusions

Instructional diagram: only shows relationships needed to understand this concept, omitting specific implementation details.

Going a Step Deeper

Common implementations split documents, build retrieval indexes, search for relevant segments, and then include the results in prompts and context. Retrieval can use keyword matching, vector similarity, or hybrid approaches. Routine RAG queries mainly change the reference information for the current request, eliminating the need to train the model for every new document.

What It Cannot Guarantee

RAG does not guarantee that the information is correct, nor does it ensure the model will strictly follow the provided materials. Its main difference from fine-tuning lies in whether knowledge is retrieved for reference or incorporated through training to adjust parameters. Practical systems can combine both approaches.

Next Steps

Sources

Information verified on 2026-09-09; the original paper is used to illustrate the mechanism, and the examples in the text are for instructional purposes.