Hướng dẫn về Retrieval-Augmented Generation (RAG)
Cách RAG hoạt động — đó retrieve-thì-generate pattern behind Google AI Overviews, ChatGPT Tìm kiếm, và Perplexity — và điều đó có nghĩa là gì cho getting nội dung của bạn cited.
Ngôn ngữ
RAG (Retrieval-Augmented Generation) là đó retrieve-thì-generate pattern behind AI tìm kiếm. Điều này chạy hai phases tại query time — retrieval (tìm relevant passages từ an external chỉ mục) và augmented generation (feed những passages để an LLM để ghi một grounded, cited câu trả lời) — không có bao giờ thay đổi đó model weights. đây là cách AI các câu trả lời cover information beyond một model training cutoff. Đó retrieval phase chains chunking → embeddings → vector tìm kiếm → re-xếp hạng → top-k passages. RAG reduces hallucinations nhưng không eliminate them — và insufficient retrieved context có thể làm them tệ hơn. Cho SEO có không tách biệt AI chỉ mục: đang crawlable, được lập chỉ mục, và structured vào clear, self-contained passages là đó prerequisite cho đang retrieved và cited.
gốc RAG architecture combined language model với information retrieved từ external chỉ mục during generation. Evidence for this claim The original RAG paper combined a pretrained sequence-to-sequence model with a non-parametric dense-vector index retrieved during generation. Scope: Lewis et al.'s 2020 RAG architecture and experiments, not every modern retrieval system. Confidence: high · Verified: Lewis et al.: Retrieval-Augmented Generation Modern nền tảng tài liệu dùng giống nhau rộng retrieve-sau đó-generate ý tưởng. Evidence for this claim Google Cloud describes RAG as retrieving relevant information from external knowledge sources and providing it to a model to improve generated responses. Scope: General RAG architecture in Google Cloud documentation; quality depends on retrieval, source quality, and generation. Confidence: high · Verified: Google Cloud: RAG overview
Tóm tắt — RAG (Retrieval-Augmented Generation) là Cách AI các công cụ tìm kiếm look điều lên trước khi họ câu trả lời. thay vì replying purely từ memory, hệ thống đầu tiên retrieves relevant passages từ tìm kiếm chỉ mục, sau đó generates câu trả lời dựa trên Điều gì nó tìm thấy. Đó là lý làm Google AI Overviews, ChatGPT Tìm kiếm, và Perplexity có thể cite fresh các trang web — và Vì sao là trong chỉ mục vẫn matters.
Điều gì RAG là
lớn language model ( LLM, điều behind ChatGPT và similar tools) learns từ huge pile của text during training. nhưng đó training có cutoff date, và model có thể’t possibly memorize mọi thứ — so on của nó own nó either không know gần đây hoặc niche facts, hoặc nó làm điều gì đó lên đó sounds right.
RAG các cách sửa đó by letting model look điều lên. Khi bạn ask câu hỏi, RAG hệ thống làm hai điều trong order:
- Retrieval — nó searches chỉ mục (như Google hoặc Bing) và pulls lại passages phần lớn relevant để của bạn câu hỏi.
- Augmented generation — nó hands những điều đó passages để LLM, mà ghi câu trả lời dựa trên them và thường hiển thị links để sources.
simplest way để picture nó: thay vì answering từ memory alone, AI làm của nó homework đầu tiên.
nhanh ví dụ
Ask an AI công cụ tìm kiếm “what changed in the latest iPhone?” (bản dịch) «điều gì changed trong đó latest iPhone?» Đó model đã không trained on một sản phẩm đó launched cuối cùng week. Với RAG, điều này searches đó web, retrieves vài gần đây các bài viết, và ghi của nó câu trả lời từ những — với citations bạn có thể nhấp. Không có RAG, điều này sẽ either chẳng hạn điều này không know hoặc guess.
Vì sao điều này quan trọng để bạn
Ở đây part đó surprises mọi người: RAG không sử dụng tách biệt “AI chỉ mục.” Google AI Overviews retrieve từ Google thông thường tìm kiếm chỉ mục. ChatGPT Tìm kiếm launched on Bing chỉ mục và cũng chạy của nó own crawler (OAI-SearchBot) — OpenAI hasn’t đã nói chính xác Cách hai là mixed hôm nay. Either way, giống nhau basics đó có luôn mattered — là crawlable, getting được lập chỉ mục, writing rõ ràng — là chính xác Điều gì decides liệu của bạn nội dung có thể là retrieved và cited trong AI câu trả lời.
khác điều để know: RAG reduces sai các câu trả lời (hallucinations) nhưng không eliminate them. AI có thể vẫn misread Điều gì nó retrieved. So là clearest, phần lớn trực tiếp nguồn on topic genuinely helps.
Muốn thực mechanics — embeddings, chunking, re-xếp hạng, naive so với. agentic RAG, và SEO playbook? Chuyển để Nâng cao tab.
Lewis và colleagues’ 2020 hệ thống paired sequence generation với dense retrieval từ non-parametric chỉ mục. Evidence for this claim The original RAG paper combined a pretrained sequence-to-sequence model with a non-parametric dense-vector index retrieved during generation. Scope: Lewis et al.'s 2020 RAG architecture and experiments, not every modern retrieval system. Confidence: high · Verified: Lewis et al.: Retrieval-Augmented Generation Google Cloud hiện tại overview defines RAG nhiều hơn broadly as supplying retrieved external knowledge để model. Evidence for this claim Google Cloud describes RAG as retrieving relevant information from external knowledge sources and providing it to a model to improve generated responses. Scope: General RAG architecture in Google Cloud documentation; quality depends on retrieval, source quality, and generation. Confidence: high · Verified: Google Cloud: RAG overview
Tóm tắt — RAG là hai-phase, inference-time pattern: retrieval (tìm relevant passages trong external corpus) sau đó augmented generation (feed những điều đó passages để LLM để produce grounded, cited câu trả lời). weights không bao giờ thay đổi — nó combines model parametric memory với non-parametric memory retrieved trực tiếp. retrieval phase chains chunking → embeddings → vector tìm kiếm → re-xếp hạng → top-k. “Naive” RAG là retrieve-sau đó-generate; Nâng cao RAG adds query rewriting và re-xếp hạng; agentic RAG adds iterative, multi-hop retrieval. Retrieval có thể ground các câu trả lời nhưng không bảo đảm correctness; trong một Gemma evaluation, insufficient context coincided với nhiều hơn không đúng các câu trả lời. Đối với SEO: có không tách biệt AI chỉ mục; crawlability, lập chỉ mục, và passage-cấp độ clarity là prerequisites cho là retrieved.
Đó hai phases (và vì sao “inference time” (bản dịch) «inference time» là đó toàn bộ point)
Five stages run left to right at inference time. Chunking splits documents into retrievable passages. Embeddings represent each passage as a dense vector. Vector search retrieves candidates and some systems combine it with BM25 keyword search. Re-ranking re-scores and narrows the candidate set. The top surviving passages enter the model context. The model's weights do not change.
© Patrick Stox LLC · CC BY 4.0 ·
Two sources feed one generation step. Parametric memory is knowledge encoded in the model weights during training and is limited by the training data and cutoff. Non-parametric memory consists of passages retrieved from an external index at query time. Generation uses both while the weights remain unchanged, producing an answer that can be grounded in and cite the retrieved sources; this does not guarantee correctness.
© Patrick Stox LLC · CC BY 4.0 ·
Break acronym apart và bạn có model: Retrieval plus Augmented Generation. query xuất hiện trong; hệ thống retrieves phần lớn relevant passages từ external corpus; nó injects những điều đó passages vào LLM context window; LLM generates câu trả lời grounded trong them.
Đó detail đó mọi người nhận sai: này happens tại inference time, và đó model weights là không bao giờ touched. RAG không phải training và điều này không phải fine-tuning. Đó original 2020 paper từ Patrick Lewis và colleagues tại Facebook AI Research được diễn đạt điều này as combining hai kinds of memory — parametric memory (knowledge baked vào đó weights during training) và non-parametric memory (knowledge retrieved trực tiếp từ an chỉ mục). RAG dùng cả hai tại khi. AWS diễn đạt đó practical case plainly: retraining một foundation model cho fresh hoặc domain-cụ thể knowledge là expensive, và “RAG is a more cost-effective approach to introducing new data to the LLM.” (bản dịch) «RAG là một hơn cost-effective approach để introducing new dữ liệu để đó LLM.»
(Đó naming, cho điều gì đây là worth, đã là an accident. Lewis sau đó admitted: “We definitely would have put more thought into the name had we known our work would become so widespread… We always planned to have a nicer sounding name, but when it came time to write the paper, no one had a better idea.” (bản dịch) «We definitely sẽ có put hơn thought vào đó name đã có we known của chúng ta hoạt động sẽ become so widespread… We luôn planned để có một nicer sounding name, nhưng khi điều này nghĩ ra time để ghi đó paper, không một đã có một tốt hơn ý tưởng.»)
Bên trong retrieval phase
“Retrieve the relevant passages” (bản dịch) «Retrieve đó relevant passages» là đang làm một lot of hoạt động trong đó sentence. Trong một real hệ thống đây là một pipeline:
- Chunking. Documents nhận split vào retrievable pieces. Chunk size là một real tradeoff — cũng nhỏ và một passage loses của nó context; cũng lớn và điều này floods đó token budget với irrelevance. Strategies range từ fixed token được tính (100/256/512) để recursive/sliding windows để “Small2Big” (retrieve một nhỏ sentence, trả về của nó parent chunk cho generation).
- Embeddings. Mỗi chunk là turned vào một dense vector — một numeric representation of của nó meaning — so similarity là computed semantically, không by từ khóa match. Này là vì sao nội dung về một topic nhận retrieved ngay cả khi điều này không dùng đó chính xác query phrasing.
- Vector tìm kiếm. Đó query là embedded cũng, và đó hệ thống tìm thấy đó chunks whose vectors sit closest để điều này. Hầu hết production stacks chạy hybrid tìm kiếm — dense vector retrieval plus BM25 từ khóa tìm kiếm — vì mỗi catches recall đó other misses.
- Re-xếp hạng. MỘT tách biệt model re-scores đó candidates by relevance để đó query và reorders them, “effectively reducing the overall document pool.” (bản dịch) «effectively reducing đó overall document pool.» Chỉ đó top survivors làm điều này vào đó context.
- Top-k vào đó prompt. Đó best passages là concatenated với người dùng query và handed để đó generator.
Chunking là đó fragile link. Anthropic identified đó “traditional RAG solutions remove context when encoding information” (bản dịch) «truyền thống RAG các giải pháp xóa context khi encoding information» — một chunk pulled out of của nó document loses đó xung quanh context đó đã làm điều này có ý nghĩa. Của họ Contextual Retrieval technique (prepending chunk-cụ thể context trước lập chỉ mục) reduced failed retrievals by 49%, và by 67% combined với re-xếp hạng. đó là một mạnh tín hiệu đó chunking vấn đề là real — và đó self-contained, context-rich passages là easier để retrieve correctly.
Naive, Nâng cao, và agentic RAG
survey literature (Gao et al., 2023) splits RAG vào hữu ích taxonomy:
- Naive RAG — “a traditional process that includes indexing, retrieval, and generation.” (bản dịch) «một truyền thống xử lý đó bao gồm lập chỉ mục, retrieval, và generation.» Retrieve top-k khi, generate khi. Điều này “struggles with precision and recall, leading to the selection of misaligned or irrelevant chunks.” (bản dịch) «struggles với precision và recall, leading để đó selection of misaligned hoặc irrelevant chunks.»
- Advanced RAG — adds “pre-retrieval and post-retrieval strategies.” (bản dịch) «pre-retrieval và post-retrieval strategies.» Pre-retrieval: query rewriting và tốt hơn lập chỉ mục (including HyDE, nơi đó model generates một hypothetical câu trả lời, embeds đó, và retrieves documents đó trông giống các câu trả lời thay vì các câu hỏi). Post-retrieval: re-xếp hạng và context compression.
- Modular / agentic RAG — đó model retrieves, reasons về điều gì là vẫn bị thiếu, và retrieves again, iterating trên multiple hops. Này là đó hiện tại state of AI tìm kiếm. As Michael King put điều này: “The retrieve-once-then- generate pattern that defined the first wave is obsolete… Agentic RAG is now the default.” (bản dịch) «Đó retrieve-khi-thì- generate pattern đó được định nghĩa đó đầu tiên wave là obsolete… Agentic RAG là hiện tại đó default.»
điều này matters Đối với SEO vì nội dung hiện tại có để survive multiple retrieval rounds và contradiction-kiểm tra — không chỉ single retrieval truyền.
Làm RAG eliminate hallucinations? Không.
Two bars report Gemma's incorrect-answer rate in one Google Research evaluation. With no context, the rate is 10.2 percent. With insufficient context, the rate is 66.1 percent. The comparison comes from Google Research's ICLR 2025 sufficient-context study and should not be generalized to every model, dataset, or retrieval system.
RAG có thể ground các câu trả lời trong retrieved sources, nhưng đó LLM có thể vẫn misread hoặc over-interpret điều gì điều này pulled. Google Research (ICLR 2025) được ghi lại một counterintuitive kết quả trong một evaluation: Gemma produced incorrect các câu trả lời on 10,2% of các câu hỏi với không context và 66,1% với insufficient context. Đó researchers báo cáo đó models có thể “excel with sufficient context but fail to recognize when context is insufficient.” (bản dịch) «excel với sufficient context nhưng fail để recognize khi context là insufficient.» Treat đó as một model- và evaluation-cụ thể warning, không proof đó retrieval universally gây ra tệ hơn các câu trả lời. Đó practical lesson là hẹp hơn: retrieval quality và context sufficiency cần để là evaluated thay vì assumed. Google operationalized đó finding as an LLM re-ranker trong của nó Vertex AI RAG Engine.
RAG so với. fine-tuning
những điều này nhận conflated constantly, và họ’re fundamentally khác:
- RAG retrieves external information tại query time. Weights unchanged. Best cho fresh/thay đổi information, citation requirements, và cost. Đó survey được tìm thấy “RAG consistently outperforms [unsupervised fine-tuning], for both existing knowledge encountered during training and entirely new knowledge.” (bản dịch) «RAG consistently outperforms [unsupervised fine-tuning], cho cả hai existing knowledge encountered during training và hoàn toàn new knowledge.»
- Fine-tuning modifies đó model weights trong một tách biệt training chạy. Best cho thay đổi style và behavior, hoặc teaching ổn định domain knowledge đó không thay đổi.
bạn’d reach cho RAG để làm model know mới nhất facts; bạn’d reach cho fine-tuning để thay đổi Cách nó talks.
RAG trong wild: Google, ChatGPT, Perplexity
- Google AI Overviews. Google calls RAG “a technique (also known as grounding)… relying on our core Search ranking systems to retrieve relevant, up-to-date web pages from our Search index.” (bản dịch) «một technique (cũng known as grounding)… relying on của chúng ta cốt lõi Tìm kiếm xếp hạng các hệ thống để retrieve relevant, lên-để-date web các trang từ của chúng ta Tìm kiếm chỉ mục.» Hai điều follow. Đầu tiên, có không tách biệt AI chỉ mục — “our generative AI features on Google Search are rooted in our core Search ranking and quality systems.” (bản dịch) «của chúng ta generative AI features on Google Search là rooted trong của chúng ta cốt lõi Tìm kiếm xếp hạng và quality các hệ thống.» Second, Google chạy query fan-out: “concurrent, related queries generated by the model to request more information.” (bản dịch) «concurrent, related các truy vấn generated by đó model để yêu cầu hơn information.» MỘT single câu hỏi có thể spawn multiple sub-các truy vấn, mỗi retrieving khác nhau nội dung — so nội dung của bạn có để satisfy đó implied sub-các câu hỏi, không chỉ đó head query.
- ChatGPT Tìm kiếm. Launched (October 2024) với Bing as của nó dữ liệu partner, và OpenAI’s own crawler tài liệu xác nhận OAI-SearchBot làm independent fetching và lập chỉ mục cho tìm kiếm citations, tách biệt từ GPTBot training-crawl. OpenAI hasn’t published đó hiện tại retrieval mix giữa Bing và của nó own chỉ mục, và OpenAI có since positioned ChatGPT Tìm kiếm as một standalone đối thủ để Bing thay vì một wrapper khoảng điều này — so treat “it’s basically Bing” (bản dịch) «đây là basically Bing» as một simplification. Đó được ghi lại, actionable lever là hẹp hơn và hơn durable: không block OAI-SearchBot trong robots.txt, vì đó là đó crawler OpenAI itself names as đó một đó indexes nội dung cho tìm kiếm citations.
- Perplexity. Được xây dựng on hybrid retrieval (Vespa.ai — BM25 + dense) với custom embedding models và một strict re-xếp hạng ngưỡng: by bên thứ ba analysis, chỉ đó top ~30% of 60-plus retrieved sources survive để đó generation stage, và “citations are not retrofitted post-generation — they are structurally assigned during context assembly.” (bản dịch) «citations không phải retrofitted post-generation — they là structurally assigned during context assembly.» Deep Research chạy đó agentic loop trên dozens of searches.
Điều gì RAG có nghĩ là Đối với SEO
Strip away jargon và playbook là concrete:
- Đang trong đó chỉ mục là đó prerequisite — đầy đủ dừng. Không tách biệt AI chỉ mục có nghĩa là đó crawl → chỉ mục → retrieve chain có để là intact. Nếu một trang không thể là được crawl và được lập chỉ mục, điều này không thể là retrieved vào an AI câu trả lời. Đó giống nhau là đúng cho đó AI engines đó xây dựng của họ own pools: AI các crawler như OAI-SearchBot và PerplexityBot có để là được phép để fetch bạn, hoặc bạn là invisible để những các câu trả lời.
- Ghi self-contained passages. RAG retrieves fragments, không toàn bộ các trang. As iPullRank Francine Monahan put điều này, AI các hệ thống examine “fragments of pages rather than the page as a whole” (bản dịch) «fragments of các trang thay vì đó trang as một toàn bộ» — so craft “stand-out passages and phrases” (bản dịch) «stand-out passages và phrases» đó câu trả lời một cụ thể câu hỏi on của họ own. Này là chính xác đó H2/H3 structure và clear topic sentences good SEO đã rewards. Google explicitly says không để chop nội dung của bạn vào tiny pieces cho AI — well-structured nội dung chunks well on của nó own.
- Cover đó sub-topics. Query fan-out có nghĩa là một câu hỏi có thể trigger nhiều retrievals. Depth trên related sub-các câu hỏi beats một trang stuffed khoảng một single từ khóa.
- Authority drives citation hơn xếp hạng position. Từ an 8 000-citation analysis: “Strong organic search presence and broad web visibility leads to AI citations, not the other way around” (bản dịch) «Mạnh organic tìm kiếm presence và rộng web visibility dẫn đến AI citations, không đó other way khoảng» — và “highly authoritative content from a lower-ranking page” (bản dịch) «highly có thẩm quyền nội dung từ một thấp hơn-xếp hạng trang» sometimes nhận cited over một ít hơn credible top-xếp hạng một. My own dữ liệu lines lên (từ my AI Overview citation research): mentions on heavily-linked các trang là đó strongest predictor of AI Overview inclusion (ρ ≈ 0,70), và branded web mentions correlated ~0,66 trên 75 000 brands.
- Fresh nội dung có an edge. AI citations skew meaningfully fresher hơn organic kết quả, so currency matters.
nếu bạn muốn một-sentence version: RAG đã không replace SEO — nó raised stakes on parts của SEO đó là luôn về là findable và là clear.
AI summary
condensed take on Nâng cao version:
- RAG = Retrieval + Augmented Generation. Hai phases tại inference time: retrieve relevant passages từ external corpus, sau đó feed them để LLM để generate grounded, cited câu trả lời. ** model weights không bao giờ thay đổi** — nó không training và không fine-tuning.
- nó combines hai memories: parametric (baked vào weights) + non-parametric (retrieved trực tiếp). đó Cách AI các câu trả lời cover information past training cutoff.
- Retrieval là pipeline: chunking → embeddings → vector tìm kiếm (thường hybrid với BM25) → re-xếp hạng → top-k passages vào prompt. Chunking là fragile link; context-rich passages retrieve tốt hơn (Anthropic cut thất bại retrievals 49%).
- Three flavors: naive (retrieve-sau khi), Nâng cao (query rewriting, HyDE, re-xếp hạng), và agentic (iterative multi-hop) — agentic là hiện tại AI-tìm kiếm default.
- nó reduces, không eliminates, hallucinations. với insufficient context, một model hallucination rate jumped 10,2% → 66,1% — bad retrieval có thể beat không retrieval.
- RAG so với. fine-tuning: RAG cho fresh/thay đổi facts + citations + cost; fine-tuning cho style/behavior và ổn định knowledge.
- Engines: Google AI Overviews retrieve từ cốt lõi chỉ mục (không tách biệt AI chỉ mục) với query fan-out; ChatGPT Tìm kiếm launched on Bing chỉ mục và cũng chạy của nó own crawler, OAI-SearchBot — chính xác hiện tại mix không phải published, so không block OAI-SearchBot; Perplexity qua hybrid retrieval với strict re-xếp hạng ngưỡng và citations assigned during context assembly.
- SEO: là crawlable + được lập chỉ mục là prerequisite; ghi self-contained passages; cover sub-topics (fan-out); authority/E-E—T drives citation nhiều hơn xếp hạng position; fresh nội dung có edge.
Tài liệu chính thức
Chính-nguồn tài liệu và definitions từ providers.
- Google Hướng dẫn để Optimizing cho Generative AI Features — defines RAG as grounding over cốt lõi Tìm kiếm chỉ mục; covers query fan-out.
- AI Overviews và AI Chế độ trong Tìm kiếm — xác nhận không additional requirements beyond tiêu chuẩn lập chỉ mục và snippet eligibility.
- RAG và grounding on Vertex AI — Google Cloud retrieve-sau đó-generate definition (Burak Gokturk).
- Deeper insights vào RAG: role của sufficient context — Google Research (ICLR 2025) on insufficient-context thất bại chế độ.
Microsoft / Azure
- RAG và generative AI — Azure AI Tìm kiếm — RAG được định nghĩa as grounding trong proprietary nội dung; query understanding, token constraints, và move để agentic retrieval.
OpenAI
- Overview của OpenAI Các crawler — xác nhận OAI-SearchBot làm independent fetching/lập chỉ mục cho ChatGPT Tìm kiếm citations, tách biệt từ GPTBot training crawl; không tiết lộ hiện tại mix với Bing chỉ mục.
Anthropic
- Introducing Contextual Retrieval — chunk-context-mất mát vấn đề và measured khắc phục (49% / 67% ít hơn thất bại retrievals).
AWS
- Điều gì là Retrieval-Augmented Generation? — sạch three-stage explainer và RAG-so với-retraining cost argument.
Foundational papers
- Retrieval-Augmented Generation cho Knowledge-Intensive NLP Tasks — Lewis et al., NeurIPS 2020 ( gốc RAG paper; parametric so với. non-parametric memory).
- Retrieval-Augmented Generation cho LLMs: Survey — Gao et al. ( naive / Nâng cao / modular taxonomy, HyDE, re-xếp hạng).
Quotes từ nguồn
On—record statements từ providers và gốc researchers. Deep links jump để quoted passage nơi khả dụng.
Google — RAG là grounding, over cốt lõi chỉ mục
- “A technique (also known as grounding) used to improve the quality, accuracy, and freshness of AI responses by relying on our core Search ranking systems to retrieve relevant, up-to-date web pages from our Search index.” (bản dịch) «MỘT technique (cũng known as grounding) được dùng để improve đó quality, độ chính xác, và freshness of AI các phản hồi by relying on của chúng ta cốt lõi Tìm kiếm xếp hạng các hệ thống để retrieve relevant, lên-để-date web các trang từ của chúng ta Tìm kiếm chỉ mục.» — Google Search Central, AI optimization hướng dẫn. Nhảy đến trích dẫn
- “Our generative AI features on Google Search are rooted in our core Search ranking and quality systems.” (bản dịch) «Của chúng ta generative AI features on Google Search là rooted trong của chúng ta cốt lõi Tìm kiếm xếp hạng và quality các hệ thống.» — Google Search Central, AI optimization hướng dẫn.
Google Cloud — retrieve-sau đó-generate definition
- “Retrieval Augmented Generation (RAG), a technique developed to mitigate these challenges, first ‘retrieves’ facts about a question, then provides those facts to the model before it ‘generates’ an answer – this is what we mean by grounding.” (bản dịch) «Retrieval Augmented Generation (RAG), một technique developed để mitigate những challenges, đầu tiên ‘retrieves’ facts về một câu hỏi, thì cung cấp những facts để đó model trước điều này ‘generates’ an câu trả lời – này là điều gì we có nghĩa là by grounding.» — Burak Gokturk, VP & GM, Cloud AI, Google Cloud (June 27, 2024). Nhảy đến trích dẫn
** gốc RAG paper — parametric so với. non-parametric memory**
- “retrieval-augmented generation (RAG) — models which combine pre-trained parametric and non-parametric memory for language generation.” (bản dịch) «retrieval-augmented generation (RAG) — models mà combine pre-trained parametric và non-parametric memory cho language generation.» — Lewis et al., NeurIPS 2020.
Patrick Lewis, lead tác giả — on name (qua NVIDIA Blog, Rick Merritt)
- “We definitely would have put more thought into the name had we known our work would become so widespread.” (bản dịch) «We definitely sẽ có put hơn thought vào đó name đã có we known của chúng ta hoạt động sẽ become so widespread.»
- “We always planned to have a nicer sounding name, but when it came time to write the paper, no one had a better idea.” (bản dịch) «We luôn planned để có một nicer sounding name, nhưng khi điều này nghĩ ra time để ghi đó paper, không một đã có một tốt hơn ý tưởng.» Đọc bài đưa tin
Microsoft — RAG as grounding trong của bạn nội dung
- “Retrieval-augmented generation (RAG) is a pattern that extends LLM capabilities by grounding responses in your proprietary content.” (bản dịch) «Retrieval-augmented generation (RAG) là một pattern đó extends LLM capabilities by grounding các phản hồi trong của bạn proprietary nội dung.» — Microsoft, Azure AI Tìm kiếm tài liệu.
Anthropic — chunking vấn đề
- “traditional RAG solutions remove context when encoding information.” (bản dịch) «truyền thống RAG các giải pháp xóa context khi encoding information.» — Anthropic, Contextual Retrieval (Sept 19, 2024). Đọc đó post
AWS — RAG so với. retraining
- “Retrieval-Augmented Generation (RAG) is the process of optimizing the output of a large language model, so it references an authoritative knowledge base outside of its training data sources before generating a response.” (bản dịch) «Retrieval-Augmented Generation (RAG) là đó xử lý of optimizing đó output of một lớn language model, so điều này references an có thẩm quyền knowledge base bên ngoài of của nó training dữ liệu sources trước generating một phản hồi.» — AWS.
- “RAG is a more cost-effective approach to introducing new data to the LLM.” (bản dịch) «RAG là một hơn cost-effective approach để introducing new dữ liệu để đó LLM.» — AWS.
OpenAI — của nó own crawler cho ChatGPT Tìm kiếm
- “OpenAI uses OAI-SearchBot and GPTBot robots.txt tags to enable webmasters to manage how their sites and content work with AI… a webmaster can allow OAI-SearchBot in order to appear in search results while disallowing GPTBot to indicate that crawled content should not be used for training.” (bản dịch) «OpenAI dùng OAI-SearchBot và GPTBot robots.txt tags để enable quản trị viên web để manage cách của họ các trang và nội dung hoạt động với AI… một quản trị viên web có thể cho phép OAI-SearchBot trong order để xuất hiện trong kết quả tìm kiếm trong khi disallowing GPTBot để indicate đó được crawl nội dung không nên là dùng cho training.» — OpenAI, Overview of OpenAI Các crawler. Đọc đó tài liệu
Michael King, iPullRank — agentic shift (công cụ tìm kiếm Land)
- “The retrieve-once-then-generate pattern that defined the first wave is obsolete… Agentic RAG is now the default.” (bản dịch) «Đó retrieve-khi-thì-generate pattern đó được định nghĩa đó đầu tiên wave là obsolete… Agentic RAG là hiện tại đó default.» Đọc bài đưa tin
RAG bảng tra nhanh
** pipeline, end để end**
query → [retrieval: chunk · embed · vector search (+BM25) · re-rank · top-k] → augment (passages into context) → generate (LLM writes grounded, cited answer)
RAG so với. fine-tuning
| RAG | Fine-tuning | |
|---|---|---|
| Thay đổi model weights? | Không | Có |
| Khi nó happens | Inference (query time) | Tách biệt training chạy |
| Best cho | Fresh/thay đổi facts, citations, cost | Style, behavior, ổn định domain knowledge |
| Cập nhật knowledge by | Re-lập chỉ mục corpus | Retraining |
** three RAG generations**
| Flavor | Điều gì nó làm | nơi bạn see nó |
|---|---|---|
| Naive | Retrieve top-k sau khi, generate sau khi | Sớm chatbots, đơn giản Q& |
| Nâng cao | + query rewriting, HyDE, re-xếp hạng, compression | phần lớn production RAG |
| Agentic | Iterative multi-hop: retrieve → reason → retrieve again | Google AI Mode, Perplexity Deep Research, ChatGPT Tìm kiếm |
Engine retrieval pools tại glance
| Engine | Retrieves từ | Note |
|---|---|---|
| Google AI Overviews | Google cốt lõi chỉ mục | Không tách biệt AI chỉ mục; query fan-out |
| ChatGPT Tìm kiếm | Bing chỉ mục + OpenAI’s own crawler | không block OAI-SearchBot; chính xác mix undisclosed |
| Perplexity | Hybrid (Vespa.ai) | Strict re-xếp hạng ngưỡng; citations assigned during assembly |
Fast facts
- RAG = Retrieval + ****ugmented Generation; coined trong Lewis et al., 2020.
- nó inference-time — weights không bao giờ thay đổi.
- Hallucination không phải solved: insufficient context took một model từ 10,2% → 66,1%.
- Context-aware chunking cut thất bại retrievals by 49% (67% với re-xếp hạng).
- không pre-”chunk” của bạn nội dung cho AI — clear H2/H3 structure chunks well on của nó own.
mental models
1. Retrieve → Augment → Generate. mỗi RAG hệ thống là những điều này three moves. Khi AI câu trả lời là sai, locate mà stage thất bại: đã làm nó retrieve right passages, đã làm nó truyền đủ context, hoặc đã làm model misgenerate từ good sources? phần lớn AI-visibility các vấn đề là retrieval các vấn đề, không generation các vấn đề.
2. Parametric so với. non-parametric memory. model có parametric knowledge (frozen trong của nó weights, capped tại của nó training cutoff) và non-parametric knowledge (retrieved trực tiếp). Xuất bản nội dung có thể’t touch weights — nhưng nó có thể feed trực tiếp retrieval. đó đểàn bộ reason SEO vẫn áp dụng để AI tìm kiếm.
3. RAG so với. fine-tuning là knowledge-so với-behavior split. Cần model để know new hoặc thay đổi facts? RAG. Cần để thay đổi Cách nó behaves hoặc ghi? Fine-tuning. không fine-tune để thêm facts đó thay đổi weekly.
4. Retrieval quality là đó bottleneck — và điều này cuts cả hai ways. Tốt hơn retrieval beats một bigger model. Và insufficient retrieval có thể là tệ hơn hơn none. So đó goal cho nội dung của bạn không chỉ “get retrieved” (bản dịch) «nhận retrieved» — đây là “get retrieved as a sufficient, self-contained passage” (bản dịch) «nhận retrieved as một sufficient, self-contained passage» đó lets đó model câu trả lời definitively.
5. crawl → chỉ mục → retrieve chain. có không tách biệt AI chỉ mục. nếu trang fails tại crawl hoặc chỉ mục, nó có thể không bao giờ reach retrieval — cho Google RAG hoặc cho AI engines building của họ own pools. khắc phục chain đầu tiên; optimize passages thứ hai.
Tự kiểm tra: Retrieval-augmented generation
các tài nguyên worth của bạn time
My related writing & research
- Điều gì chúng ta thực ra Know về Optimizing cho LLM Tìm kiếm — Ahrefs’ ghi-lên sử dụng my dữ liệu: mentions on heavily-linked các trang là strongest predictor của AI Overview inclusion (ρ ≈ 0,70).
- Generative Engine Optimization — SEO phản hồi để RAG-powered tìm kiếm landscape.
- GEO? AEO? LLMO? Điều gì với All điều này AI SEO Stuff? — my Ahrefs Evolve 2025 talk on AI tìm kiếm landscape và Vì sao lập chỉ mục prerequisite hasn’t changed.
** foundational papers**
- Retrieval-Augmented Generation cho Knowledge-Intensive NLP Tasks — Lewis et al., 2020 ( origin).
- RAG cho LLMs: Survey — Gao et al. ( naive/Nâng cao/modular taxonomy).
từ others
- Cách AI các công cụ tìm kiếm Hoạt động — Ryan Law (Ahrefs) on RAG as grounding mechanism.
- Google AI Overviews: All bạn cần Know — Ong & Law (Ahrefs) on RAG over cốt lõi chỉ mục.
- Điều gì là Retrieval-Augmented Generation? — NVIDIA (bao gồm Lewis naming anecdote).
- Cách Retrieval-Augmented Generation là Redefining SEO — Francine Monahan, iPullRank (passage-cấp độ optimization).
- Beyond RAG: Vì sao mỗi AI tìm kiếm nền tảng là hiện tại agentic — Michael King, công cụ tìm kiếm Land.
- Cách Perplexity AI Các câu trả lời Hoạt động — Ishtiaque Ahmed, kỹ thuật breakdown của retrieval/xếp hạng/citation pipeline.
- Cách nhận cited by AI: SEO insights từ 8 000 AI citations — James Allen, công cụ tìm kiếm Land; authority và E-E—T drive AI citations nhiều hơn xếp hạng position.
- Cách Perplexity dùng Vespa.ai — Vespa.ai đầu tiên-party account của Perplexity hybrid BM25 + dense retrieval architecture.
- Retrieval-augmented generation — Wikipedia — hữu ích reference overview; covers RAG poisoning và hallucination caveat.
Số liệu worth citing
- 10,2% → 66,1% hallucination jump — một model hallucination rate với insufficient retrieved context so với. không context tại all; bad retrieval có thể beat không retrieval. Google Research, ICLR 2025. Nguồn
- 49% ít hơn failed retrievals từ context-aware chunking (Contextual Embeddings), rising để 67% khi combined với re-xếp hạng. Anthropic, 2024. Nguồn
- ρ ≈ 0,70 — mentions on heavily-linked các trang là đó strongest predictor of Google AI Overview inclusion trong my research; branded web mentions correlated ~0,66 trên 75 000 brands. Nguồn
- ~30% survival rate — by bên thứ ba analysis, chỉ khoảng đó top 30% of 60+ retrieved sources clear Perplexity re-xếp hạng ngưỡng vào đó generation stage. Nguồn
- RAG > unsupervised fine-tuning cho knowledge tasks — “for both existing knowledge encountered during training and entirely new knowledge.” (bản dịch) «cho cả hai existing knowledge encountered during training và hoàn toàn new knowledge.» Nguồn
Nhật ký thay đổi
Đã cập nhật 8 thg 8, 2026.
Tóm tắt biên tập và chi tiết thay đổi đã ghi nhận.Chi tiết thay đổi
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
Không thể so sánh đầy đủ — không có bản lưu trước đó cho lần sửa đổi này.
Đã cập nhật 19 thg 7, 2026.
Tóm tắt biên tập và chi tiết thay đổi đã ghi nhận.Chi tiết thay đổi
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
Không thể so sánh đầy đủ — không có bản lưu trước đó cho lần sửa đổi này.