Hướng dẫn về Chunking
Cách AI các hệ thống split của bạn các trang vào passages to embed, chỉ mục, và retrieve — chunk size, overlap, semantic so với. fixed-size, và cách điều này ties to Google passage xếp hạng.
Ngôn ngữ
Chunking là đó preprocessing step nơi AI các hệ thống split một document vào nhỏ hơn passages trước embedding và retrieving them. Đó chunk — không đó trang — là đó unit of retrieval trong AI tìm kiếm, so đang được lập chỉ mục không đủ; bạn cần một passage đó các câu trả lời một cụ thể query trong isolation. Chunk size là một trade-off (nhỏ = precise nhưng thin on context; lớn = context-rich nhưng noisy) với không size đó bảo đảm một citation, overlap ngăn boundary mất mát nhưng duplicates tokens, contextual prefixes có thể disambiguate một chunk nhưng mislead retrieval nếu họ là stale, và 'Lost trong đó Middle' — một position effect measured on cụ thể 2023 models và tasks, không mỗi model — có nghĩa là LLMs thường dùng đó bắt đầu và end of context best. Bạn không thể optimize cho một cụ thể chunk size — Google explicitly says bạn không cần to chop nội dung vào pieces — nhưng câu trả lời-đầu tiên, self-contained sections under một clear heading hierarchy chunk và cite cleanly. đây là đó giống nhau sub-document principle as Google passage xếp hạng, applied to RAG.
Tóm tắt — Chunking là Cách AI các công cụ tìm kiếm cut của bạn trang vào nhỏ hơn pieces trước khi họ store và tìm kiếm nó. họ không match query so với của bạn toàn bộ trang — họ match nó so với single passage đó các câu trả lời nó best. So getting được lập chỉ mục không phải đủ; bạn cần section đó làm sense on của nó own và các câu trả lời câu hỏi rõ ràng.
Điều gì chunking là
Chunking divides documents vào nhỏ hơn units cho embedding và retrieval các hệ thống. Evidence for this claim Retrieval systems can split files into chunks that are embedded and indexed for later search. Scope: OpenAI's retrieval implementation; chunk sizes, overlap, and indexing behavior are implementation-dependent. Confidence: high · Verified: OpenAI: Retrieval guide Chunk size và overlap là implementation choices whose best các giá trị phụ thuộc on nội dung, model, và evaluation task. Evidence for this claim Retrieval-augmented generation combines a generator with retrieved external passages or documents. Scope: The original RAG research architecture; it does not establish one optimal chunking strategy for every production system. Confidence: high · Verified: Lewis et al.: Retrieval-Augmented Generation
Khi bạn tìm kiếm Google old way, trang là unit: trang ranks, và bạn click qua. AI các hệ thống hoạt động differently. trước khi ChatGPT, Perplexity, hoặc Google AI Overviews có thể sử dụng của bạn nội dung, họ break nó vào nhỏ hơn segments — được gọi là chunks hoặc passages — và store mỗi một riêng. Khi ai đó asks câu hỏi, hệ thống goes và tìm thấy chunks đó best match nó, sau đó ghi câu trả lời từ những điều đó.
So unit của retrieval không phải của bạn trang. nó passage từ của bạn trang.
Vì sao họ bother
Hai reasons:
- Size limits. Đó models đó turn text vào searchable math (see embeddings) có thể chỉ take so nhiều text tại khi. MỘT 5 000-word trang sẽ không fit, so điều này nhận split.
- Precision. MỘT toàn bộ trang về “technical SEO” (bản dịch) «SEO kỹ thuật» là một blurry match cho một cụ thể câu hỏi like “what is crawl budget.” (bản dịch) «điều gì là ngân sách crawl.» MỘT focused 400-word section đó là chỉ về ngân sách crawl là một sharp match. Nhỏ hơn pieces let đó hệ thống tìm đó needle thay vì handing over đó haystack.
Điều này có nghĩ là gì cho bạn
Bạn có thể’t control Cách bất kỳ AI hệ thống chops up của bạn nội dung — và bạn không cần để. Điều gì Bạn có thể làm là ghi so đó mỗi section survives là pulled out on của nó own:
- Put câu trả lời đầu tiên. Lead section với definition hoặc Điểm mấu chốt claim, không three paragraphs của throat-clearing.
- giữ sections focused. Một topic hoặc câu hỏi per heading.
- Làm mỗi section self-contained. Ask yourself: nếu ai đó đọc chỉ điều này paragraph, với không có gì khoảng nó, sẽ nó vẫn làm sense?
Good news: này là chỉ clear writing. Google says outright đó bạn không cần to break nội dung của bạn vào tiny pieces cho AI — của nó các hệ thống làm đó splitting. Đó advanced version nhận vào chunk size, overlap, đó research, và cách này all connects to Google “passage ranking.” (bản dịch) «passage xếp hạng.»
TL;DR — Chunking là đó preprocessing step đó splits một document vào passages trước họ là embedded, được lập chỉ mục, và retrieved. Điều này tồn tại vì embedding models và context windows có token limits, và vì passage-level retrieval là hơn precise hơn trang-level. Đó chunk, không đó trang, là đó unit of retrieval. Chunk size là một trade-off (nhỏ = precise/thin; lớn = rich/noisy) và không size, overlap giá trị, hoặc splitter bảo đảm một citation; overlap guards boundaries nhưng costs chỉ mục size và duplication; semantic chunking không reliably tốt hơn hơn fixed-size; contextual prefixes có thể disambiguate một chunk nhưng mislead retrieval khi stale. “Lost in the Middle” (bản dịch) «Lost trong đó Middle» — một position effect được ghi lại on cụ thể 2023 models và tasks — có nghĩa là an LLM thường dùng đó bắt đầu và end of của nó context best. Bạn không thể optimize cho một cụ thể chunk size — và Google says bạn không nên try — nhưng câu trả lời-đầu tiên, self-contained sections under một clear heading hierarchy chunk và cite cleanly. đây là đó giống nhau sub-document principle as Google passage xếp hạng, applied to RAG.
Điều gì chunking là, và Vì sao nó tồn tại
Retrieval các hệ thống operate over được lập chỉ mục units, nhưng công khai các công cụ tìm kiếm không expose universal publisher-controlled chunk size. Evidence for this claim Retrieval systems can split files into chunks that are embedded and indexed for later search. Scope: OpenAI's retrieval implementation; chunk sizes, overlap, and indexing behavior are implementation-dependent. Confidence: high · Verified: OpenAI: Retrieval guide RAG research hỗ trợ retrieve-sau đó-generate patterns không có proving đó mỗi AI tìm kiếm sản phẩm dùng giống hệt mechanics. Evidence for this claim Retrieval-augmented generation combines a generator with retrieved external passages or documents. Scope: The original RAG research architecture; it does not establish one optimal chunking strategy for every production system. Confidence: high · Verified: Lewis et al.: Retrieval-Augmented Generation
Chunking là xử lý của dividing document vào nhỏ hơn, discrete segments — chunks hoặc passages — trước khi những điều đó segments là turned vào embeddings, stored trong vector chỉ mục, và retrieved để câu trả lời queries. nó foundational step trong RAG (Retrieval-Augmented Generation) pipelines, và nó part của AI tìm kiếm SEOs underrate phần lớn, vì nó invisible: nó happens during ingestion, không tại query time.
Hai các vấn đề force nó:
- Đó token-limit vấn đề. Embedding models accept một bounded number of tokens
per input — Microsoft notes đó
text-embedding-3-smallmodel caps tại 8 191 tokens, và other models là far nhỏ hơn. LLMs có finite context windows cũng. MỘT dài trang đơn giản sẽ không fit as một single unit, so điều này nhận split. - Đó retrieval-precision vấn đề. MỘT 5 000-word trang về “technical SEO” (bản dịch) «SEO kỹ thuật» collapsed vào một vector là một fuzzy, averaged tín hiệu. MỘT 400-word section đó là cụ thể về ngân sách crawl, embedded as của nó own vector, là một sharp một. Sub-document granularity là điều gì lets một hệ thống pull đó một relevant passage out of một dài, multi-topic trang.
consequence là single phần lớn quan trọng idea ở đây: ** chunk, không trang, là unit của retrieval trong AI tìm kiếm.** là được crawl và được lập chỉ mục là necessary nhưng chư đủ. bạn cần chunk đó các câu trả lời cụ thể query, rõ ràng, on của nó own. trang xếp hạng #15 organically có thể earn AI citation trong khi #1 kết quả nhận skipped — nếu #15 trang có nhiều hơn extractable passage.
Cách chunking hoạt động, end để end
A four-stage flow begins with one long document containing several topics. The system splits and embeds focused passages as separate vectors. A query retrieves one or more best-matching passages, and those selected passages enter the model context for answer generation.
© Patrick Stox LLC · CC BY 4.0 ·
- Ingestion + splitting. AI crawler downloads trang; chunking algorithm divides text vào (thường overlapping) segments.
- Embedding. mỗi chunk becomes numerical vector và là stored trong vector chỉ mục. (đó embeddings step.)
- Retrieval + generation. query là embedded, nearest chunk vectors là pulled qua vector tìm kiếm, và retrieved chunks là đã truyền để LLM đó ghi câu trả lời với citations.
điều này toàn bộ architecture traces back để Dense Passage Retrieval (Karpukhin et al., 2020), mà showed đó retrieving by dense vector similarity beat old từ khóa approach (BM25) by 9–19% absolute on top-20 passage độ chính xác — và đó passage-level retrieval hoạt động tốt hơn hơn document-level cho answering cụ thể các câu hỏi. RAG paper (Lewis et al., 2020) named pattern và được sử dụng 100-word passages từ Wikipedia as của nó chunks. mỗi AI tìm kiếm hệ thống retrieving web nội dung là đang chạy variant của đó pipeline.
Chunking strategies
có không single algorithm — các hệ thống pick từ menu, và họ thay đổi nó theo thời gian:
- Fixed-size — split by một token hoặc character count, với some overlap. Hầu hết phổ biến. Microsoft ví dụ: “a fixed size sufficient for semantically meaningful paragraphs (for example, 200 words or 600 characters)” (bản dịch) «một fixed size sufficient cho semantically có ý nghĩa paragraphs (ví dụ, 200 words hoặc 600 characters)» với 10–15% overlap.
- Sentence / paragraph — split on natural language boundaries thay vì an arbitrary token count, bảo toàn semantic units.
- Semantic / nội dung-aware — group sentences by embedding similarity và split nơi đó topic shifts, aiming to giữ mỗi chunk về một điều.
- Hierarchical / recursive (RAPTOR) — xây dựng một tree of summaries (document → section → paragraph) so một query có thể là answered tại đó right level of abstraction. RAPTOR (Sarthi et al., 2024) reported một 20% absolute độ chính xác gain on một hard QA benchmark khi paired với GPT-4.
- Sliding window với overlap — mỗi chunk shares some tokens với của nó neighbors so một sentence split across một boundary không lost.
- Adaptive / query-dependent (Mix-of-Granularity) — một trained router picks đó chunk size per query. Đó hầu hết sophisticated approach; chưa tiêu chuẩn trong commercial các hệ thống.
Chunk size và overlap — core trade-off
Đây là lever mọi người asks về, và honest câu trả lời là nó phụ thuộc:
- Nhỏ chunks (128–256 tokens): nhiều hơn precise retrieval, nhưng họ có thể lose xung quanh context câu trả lời cần.
- Lớn chunks (512–1 024 tokens): nhiều hơn context được bảo toàn, nhưng noisier retrieval — bạn drag trong irrelevant material với relevant bit.
Đó research không crown một winner. LlamaIndex evaluation được tìm thấy 1 024 tokens optimal trong của họ setup; Chroma benchmarks được tìm thấy một 200-token recursive splitter performed consistently across các chỉ số. Ravi Theja takeaway là đó right mindset: “Identifying the best chunk size for a RAG system is as much about intuition as it is empirical evidence.” (bản dịch) «Identifying đó best chunk size cho một RAG hệ thống là as nhiều về intuition as điều này là empirical evidence.» Những numbers mô tả điều gì won trong một team benchmark, on một document set, với một embedding model và evaluation task — không một universal setting. Không chunk size, overlap giá trị, hoặc splitter bảo đảm retrieval, citation, xếp hạng, hoặc inclusion trong an câu trả lời từ an external AI hệ thống; mỗi provider pipeline picks của nó own defaults và có thể thay đổi them không có notice.
Overlap là đó underrated half of này. Không có điều này, nội dung near một boundary nhận split và lost. Microsoft khuyến nghị starting tại 25% overlap so có “smoother transitions between chunks without excessive duplication” (bản dịch) «smoother transitions giữa chunks không có excessive duplication»; other sources suggest 10–15%. Overlap không free, though: overlapping tokens nhận embedded và stored twice, mà inflates đó chỉ mục và có thể surface near-duplicate passages side by side trong một retrieved set — so đây là một trade so với chỉ mục size và redundancy, không một costless safety net. Weigh điều này so với cách thường một boundary thực ra costs bạn một missed fact, không on principle. Đó practical point cho nội dung: không assume đó hệ thống sẽ giữ một key fact intact nếu bạn bury điều này chính xác nơi một chunk có khả năng break.
So, làm semantic chunking luôn win? Không — đó là đó inconvenient finding. Vectara 2024 study được tìm thấy “performance differences are minimal” (bản dịch) «performance differences là minimal» on thực tế documents, và — crucially — đó embedding model quality mattered hơn đó chunking strategy. Khi GPT-4o generated đó các câu trả lời, đó differences giữa strategies đã là “negligible.” Đó kết quả là cụ thể to Vectara document set, models, và evaluation phương thức — đây là evidence semantic chunking không một reliable default win, không proof điều này không bao giờ helps trong any pipeline. Translation cho SEOs: obsessing over an chính xác structure matters far ít hơn hơn nội dung quality và semantic clarity.
Chunk context: Điều gì prefixes có thể (và có thể’t) khắc phục
chunk đó đọc rõ ràng để human có thể vẫn retrieve badly sau khi nó separated từ document khoảng nó — section nó sits under, entity nó thực ra về, qualifier stated hai headings up. Một được ghi lại mitigation là prepending ngắn, chunk-cụ thể context string ( document tiêu đề, section nó belongs để, Điều gì chunk là thực ra về) trước khi chunk là embedded — approach Anthropic calls contextual retrieval. Các tiêu đề, section ancestry, và contextual prefix có thể disambiguate nếu không-orphaned chunk và cut down on retrieval misses gây ra by missing context.
đó khắc phục có thất bại chế độ của của nó own, though: stale hoặc sai context không chỉ fail để help — nó actively misleads retriever toward sai chunk. prefix generated từ heading đó không lâu hơn matches section sau khi trang reorg, hoặc contextual summary đó misstates Điều gì chunk covers, là tệ hơn hơn không prefix tại all. Context preservation là pipeline design lựa chọn với của nó own lỗi chế độ, không một-way improvement Bạn có thể bolt on và forget.
Google passage xếp hạng — SEO ancestor của chunking
Google đã được đang làm sub-document granularity since dài trước “RAG” đã là một buzzword. Tại Tìm kiếm On 2020, Prabhakar Raghavan announced passage xếp hạng: “By better understanding the relevancy of specific passages, not just the overall page, we can find that needle-in-a-haystack information you’re looking for.” (bản dịch) «By tốt hơn understanding đó relevancy of cụ thể passages, không chỉ đó overall trang, we có thể tìm đó needle-trong-một-haystack information bạn là looking cho.» Điều này went trực tiếp trong U.S. English on February 10, 2021 và ảnh hưởng roughly 7% of queries.
Hai điều mọi người nhận sai về nó:
- đây là “passage ranking,” (bản dịch) «passage xếp hạng,» không “passage indexing.” (bản dịch) «passage lập chỉ mục.» Google đầu tiên announcement dùng “lập chỉ mục,” thì quickly corrected điều này: “this change doesn’t mean we’re indexing individual passages independently of pages.” (bản dịch) «này thay đổi không có nghĩa là chúng ta là lập chỉ mục individual passages independently of các trang.» Đó trang là vẫn được lập chỉ mục as một toàn bộ; đó relevant passage là an additional tín hiệu xếp hạng.
- Đó trang ranks, không đó passage. John Mueller: “Passage ranking is not about ranking a specific passage but understanding the content on a really long, not SEO optimized page, and ranking that page (not the passage) for a query where the passage is relevant.” (bản dịch) «Passage xếp hạng không phải về xếp hạng một cụ thể passage nhưng understanding đó nội dung on một thực sự dài, không SEO optimized trang, và xếp hạng đó trang (không đó passage) cho một query nơi đó passage là relevant.»
So passage xếp hạng và RAG chunking share principle — paragraph, không trang, là thường right unit để match so với cụ thể query — nhưng outcome differs: passage xếp hạng lifts trang xếp hạng; RAG chunking retrieves chunk để feed generated câu trả lời. giống nhau idea, khác machinery. ( kỹ thuật underpinning, per Dawn Anderson reporting, là DeepCT — BERT-derived contextual term weights replacing TF-IDF, so term frequency không lâu hơn equals term relevance.)
”Lost in the Middle” (bản dịch) «Lost trong đó Middle» — nơi của bạn câu trả lời sits matters
Even sau của bạn chunk là retrieved, nơi điều này lands trong đó LLM context ảnh hưởng liệu đó model thực ra dùng điều này. Đó Stanford “Lost in the Middle” (bản dịch) «Lost trong đó Middle» study (Liu et al., 2023) được tìm thấy đó “performance is often highest when relevant information occurs at the beginning or end of the input context, and significantly degrades when models must access relevant information in the middle of long contexts.” (bản dịch) «performance là thường highest khi relevant information occurs tại đó beginning hoặc end of đó input context, và significantly degrades khi models phải access relevant information trong đó middle of dài contexts.»
đó finding không apply universally theo mặc định — nó position effect measured on named multi-document QA và mấu chốt-giá trị retrieval tasks, on cụ thể model generation Liu et al. tested trong 2023. nó không proof đó mỗi hiện tại model bỏ qua evidence placed centrally trong của nó context; khác architectures, lâu hơn effective context windows, và newer training có thể hẹp hoặc widen effect. Treat nó as được ghi lại risk để design khoảng, không fixed law của mỗi model bạn’ll bao giờ là retrieved vào.
nội dung implication là vẫn concrete và thấp-risk either way: lead với của bạn câu trả lời. Put definition, Điểm mấu chốt finding, trực tiếp câu trả lời trong đầu tiên sentence của mỗi section — không buried trong paragraph four. Đây là giống nhau câu trả lời-đầu tiên (BLUF) discipline đó phục vụ human skimmers; nó chỉ happens để cũng hedge so với là dropped vào middle của context window on models nơi effect holds.
practical constraints Bạn có thể’t see
- Chrome ~30-passage limit. Dan Petrovic research suggests Chrome DocumentChunker analyzes nội dung trong ~200-word passages và “only ever considers the first 30 passages of a page.” (bản dịch) «chỉ bao giờ considers đó đầu tiên 30 passages of một trang.» Điều này tree-walks đó semantic HTML top to bottom. Implication: của bạn hầu hết quan trọng nội dung nên xuất hiện sớm — không tại đó bottom of một 10 000-word trang.
- Bạn không control đó chunker. Despina Gavoyannis (Ahrefs) là blunt: “You can’t control how Google, ChatGPT, or Perplexity chunk your content. Their pipelines change based on cost, model, and context.” (bản dịch) «Bạn không thể control cách Google, ChatGPT, hoặc Perplexity chunk nội dung của bạn. Của họ pipelines thay đổi dựa trên cost, model, và context.» Và: “Manual ‘chunk optimization’ is impossible in practice.” (bản dịch) «Manual ‘chunk optimization’ là không thể trong thực tế.»
Điều gì “chunk optimization” (bản dịch) «chunk optimization» thực ra là (và Google caveat)
Ở đây đó nuance đó hype skips. Google 2026 AI optimization hướng dẫn says plainly: “There’s no requirement to break your content into tiny pieces for AI to better understand it.” (bản dịch) «có không requirement to break nội dung của bạn vào tiny pieces cho AI to tốt hơn understand điều này.» Của nó các hệ thống “are able to understand the nuance of multiple topics on a page and show the relevant piece to users.” (bản dịch) «là able to understand đó nuance of multiple topics on một trang và cho thấy đó relevant piece to người dùng.» So không — làm không rewrite nội dung của bạn vào rigid 300-word blocks.
Nhưng đó không có nghĩa là structure là irrelevant. As Gavoyannis puts điều này, “Most SEOs using the term [chunk optimization] are just talking about good content structure.” (bản dịch) «Hầu hết SEOs dùng đó term [chunk optimization] là chỉ talking về good nội dung structure.» Đó advice underneath đó buzzword là sound; đó framing chỉ inflates của nó novelty. Duane Forrester line captures đó shift well: “If traditional SEO optimized for clicks, GenAI systems optimize for chunks… Structure still wins.” (bản dịch) «Nếu truyền thống SEO optimized cho clicks, GenAI các hệ thống optimize cho chunks… Structure vẫn wins.»
So Điều gì làm bạn thực ra làm? Ghi chunk-ready nội dung:
- Một topic per section. focused H2/H3 maps cleanly để coherent chunk.
- Câu trả lời-đầu tiên. Lead với claim; hỗ trợ nó dưới.
- Self-contained sections. Kiểm thử: sẽ điều này paragraph làm sense nếu nó appeared trong isolation? nếu không, nó sẽ không survive là chunked out.
- Appropriate length. 200–500 words per major section lines up naturally với 256–512-token chunks — hoàn tất đủ để là hữu ích, không artificially ngắn.
- Structured formats. Các bảng và lists cho các hệ thống rõ ràng boundaries để detect. (Onely research: các bảng increase citation rates ~2,5x.)
Mike King framing là đó right reassurance: “chunking and writing for users is not mutually exclusive.” (bản dịch) «chunking và writing cho người dùng không phải mutually exclusive.» Đó structure đó helps một reader skim là đó structure đó chunks cleanly. bạn là không optimizing cho một robot tại đó expense of một human — đây là đó giống nhau nội dung.
Passage xếp hạng so với. RAG chunking — side by side
| Google passage xếp hạng | RAG chunking | |
|---|---|---|
| Điều gì nó là | xếp hạng tín hiệu | preprocessing step |
| Granularity | Passage trong trang | Chunk split trước khi embedding |
| Outcome | trang ranks cao hơn | chunk là retrieved vào câu trả lời |
| nơi nó chạy | Tại xếp hạng time | Tại ingestion (sau đó retrieval) |
| bạn control split? | Không | Không |
| Shared principle | Sub-document granularity — paragraph, không trang, thường matches cụ thể query best |
nơi điều này sits trong pipeline
Chunking là đầu tiên move trong AI retrieval: chunk → embed → store → retrieve → generate. nó feeds embeddings (mỗi chunk becomes vector), mà feed vector tìm kiếm (query matches nearest chunks), mà feeds RAG (retrieved chunks become câu trả lời). Upstream, AI các crawler là Cách của bạn nội dung nhận ingested trong đầu tiên place. cho truyền thống-tìm kiếm version của điều này toàn bộ pipeline, see Cách Tìm kiếm Hoạt động.
AI summary
condensed take on Nâng cao version:
- Chunking = splitting một document vào passages trước họ là embedded, được lập chỉ mục, và retrieved. đây là đó đầu tiên step trong một RAG pipeline và điều này happens tại ingestion, không query time.
- Đó chunk, không đó trang, là đó unit of retrieval. Đang được lập chỉ mục không đủ; bạn cần một passage đó các câu trả lời một cụ thể query on của nó own. MỘT #15 trang có thể nhận cited over một #1 nếu của nó passage là hơn extractable.
- Vì sao điều này tồn tại: embedding models và context windows có token limits, và passage-level retrieval là hơn precise hơn trang-level.
- Chunk size là một trade-off: nhỏ (128–256 tokens) = precise nhưng thin on context; lớn (512–1 024) = rich nhưng noisy. Không universal best — LlamaIndex được tìm thấy 1 024 optimal, Chroma được tìm thấy 200 — và không chunk size hoặc splitter bảo đảm một citation. Overlap (10–25%) guards chunk boundaries nhưng duplicates tokens và inflates đó chỉ mục; weigh điều này so với thực tế boundary losses.
- Semantic chunking không reliably tốt hơn hơn fixed-size (Vectara 2024, on đó study models/documents) — embedding-model quality matters hơn đó chunking strategy.
- Contextual prefixes (các tiêu đề, section ancestry — Anthropic “contextual retrieval” (bản dịch) «contextual retrieval») có thể disambiguate an orphaned chunk, nhưng một stale hoặc wrong prefix actively misleads retrieval; đây là một design lựa chọn với của nó own chế độ lỗi.
- “Lost in the Middle” (bản dịch) «Lost trong đó Middle» (Liu et al., 2023): một position effect measured on named tasks và 2023-era models, không proof mỗi hiện tại model bỏ qua central context — LLMs trong đó study dùng đó bắt đầu và end of context best, so lead mỗi section với đó câu trả lời regardless.
- Google passage xếp hạng (2020/2021, ~7% of queries) là đó giống nhau sub-document principle — nhưng đây là một tín hiệu đó ranks đó trang, không tách biệt lập chỉ mục of passages. RAG chunking retrieves một chunk.
- Bạn không thể optimize cho một chunk size, và Google says bạn không cần to chop nội dung vào tiny pieces. Bạn có thể ghi câu trả lời-đầu tiên, self-contained, focused sections under clear headings — mà là chỉ good structure.
Tài liệu chính thức
Chính-nguồn tài liệu on passage-level retrieval và chunking.
- MỘT Hướng dẫn to Google Search Xếp hạng Các hệ thống — defines passage xếp hạng as “an AI system we use to identify individual sections or ‘passages’ of a web page.” (bản dịch) «an AI hệ thống we dùng to identify individual sections hoặc ‘passages’ of một web trang.»
- Optimizing của bạn website cho generative AI features — Google stance đó bạn không cần to break nội dung vào tiny pieces (cuối cùng đã cập nhật June 15, 2026).
- Trong-Depth Hướng dẫn to Cách Google Search Hoạt động — đó crawl → chỉ mục → serve pipeline này AI version extends.
Microsoft / Azure AI Tìm kiếm
- Chunk lớn documents cho vector tìm kiếm — phần lớn thorough chính thức chunking hướng dẫn của bất kỳ major provider: chunk-size defaults (512 tokens), overlap (25% bắt đầu point), và fixed/variable/semantic technique bảng (cuối cùng đã cập nhật June 8, 2026).
OpenSearch
- Text chunking — chunking as được xây dựng-trong vector-tìm kiếm ingestion pipeline feature.
Practitioner reference (vendor tài liệu)
- Pinecone — Chunking Strategies — đó canonical strategy taxonomy, với đó kiểm thử đó anchors điều này all: “If the chunk of text makes sense without the surrounding context to a human, it will make sense to the language model as well.” (bản dịch) «Nếu đó chunk of text làm sense không có đó xung quanh context to một human, điều này sẽ làm sense to đó language model as well.»
Quotes từ nguồn
On—record statements từ Google. mỗi link là deep link đó jumps để quoted passage on nguồn trang.
Google — Điều gì passage xếp hạng là (và không phải)
- “By better understanding the relevancy of specific passages, not just the overall page, we can find that needle-in-a-haystack information you’re looking for.” (bản dịch) «By tốt hơn understanding đó relevancy of cụ thể passages, không chỉ đó overall trang, we có thể tìm đó needle-trong-một-haystack information bạn là looking cho.» — Prabhakar Raghavan, Google SVP, October 2020 (qua Search Engine Land). Jump to quote
- “this change doesn’t mean we’re indexing individual passages independently of pages.” (bản dịch) «này thay đổi không có nghĩa là chúng ta là lập chỉ mục individual passages independently of các trang.» — Google, October 20, 2020 clarification (qua Search Engine Land). Jump to quote
- “passage ranking launched yesterday afternoon Pacific Time for queries in the US in English.” (bản dịch) «passage xếp hạng launched yesterday afternoon Pacific Time cho queries trong đó US trong English.» — @searchliaison, February 11, 2021 (qua Search Engine Land). Đọc đó coverage
Google — passage xếp hạng ranks trang, không passage
- “Passage ranking is not about ranking a specific passage but understanding the content on a really long, not SEO optimized page, and ranking that page (not the passage) for a query where the passage is relevant.” (bản dịch) «Passage xếp hạng không phải về xếp hạng một cụ thể passage nhưng understanding đó nội dung on một thực sự dài, không SEO optimized trang, và xếp hạng đó trang (không đó passage) cho một query nơi đó passage là relevant.» — John Mueller, Google Search Advocate (qua Công cụ tìm kiếm Roundtable). Đọc đó coverage
Google — on liệu bạn nên chunk của bạn nội dung
- “There’s no requirement to break your content into tiny pieces for AI to better understand it.” (bản dịch) «có không requirement to break nội dung của bạn vào tiny pieces cho AI to tốt hơn understand điều này.» — Google Search Central, “Optimizing your website for generative AI features.” (bản dịch) «Optimizing của bạn website cho generative AI features.» Đọc đó hướng dẫn
Microsoft Azure AI Tìm kiếm — Vì sao chunking là necessary
- “Partitioning large documents into smaller chunks can help you stay under the maximum token input limits of chat completion and embedding models.” (bản dịch) «Partitioning lớn documents vào nhỏ hơn chunks có thể help bạn stay under đó maximum token input limits of chat completion và embedding models.» — Microsoft Azure AI Tìm kiếm tài liệu. Đọc đó tài liệu
Chunking — bảng tra nhanh
** chunking strategies, compared**
| Strategy | Cách nó splits | Strength | Watch-out |
|---|---|---|---|
| Fixed-size | Token/char count (e.g. 512 tokens) | đơn giản, fast, predictable | Cuts mid-thought không có overlap |
| Sentence / paragraph | Natural language boundaries | Preserves semantic units | Variable, uneven chunk sizes |
| Semantic | Splits nơi topic shifts (embedding similarity) | Coherent chunks | Costly; không reliably tốt hơn (Vectara 2024) |
| Hierarchical (RAPTOR) | Tree của summaries: doc → section → paragraph | Các câu trả lời tại right abstraction level | Phức tạp để xây dựng |
| Sliding window | Overlapping fixed chunks | Guards boundary context | Some duplication |
| Adaptive (Mix-của-Granularity) | Router picks size per query | phần lớn flexible | không tiêu chuẩn trong production |
Chunk size, tại glance
| Size | Tokens | Behavior |
|---|---|---|
| Nhỏ | 128–256 | Precise retrieval, thin on context |
| Medium | 512 | phổ biến default (Microsoft starting point) |
| Lớn | 1 024 | Context-rich, noisier (LlamaIndex optimum trong kiểm thử) |
| Overlap | 10–25% | Ngăn boundary mất mát; Microsoft bắt đầu tại 25% |
Fast facts
- Đó chunk, không đó trang, là đó unit of retrieval trong AI tìm kiếm.
- Embedding model cap ví dụ:
text-embedding-3-small= 8 191 tokens. - Không chunk size, overlap giá trị, hoặc splitter bảo đảm retrieval, citation, hoặc inclusion trong an AI câu trả lời — những là pipeline defaults, không universal settings.
- Overlap không free: overlapping tokens là stored twice, inflating đó chỉ mục và risking near-duplicate retrieved passages.
- Contextual prefixes (tiêu đề, section ancestry) có thể disambiguate một chunk — nhưng một stale hoặc wrong prefix misleads retrieval thay vì helping điều này.
- “Lost in the Middle” (bản dịch) «Lost trong đó Middle»: trong đó 2023 study tasks và models, LLMs dùng đó bắt đầu và end of context best → lead với đó câu trả lời regardless of model.
- Chrome DocumentChunker reportedly considers chỉ đó đầu tiên ~30 passages (~200 words mỗi) → put key nội dung sớm.
- Google: passage xếp hạng ≠ passage lập chỉ mục. Đó trang ranks; đó passage là một tín hiệu.
- Google: bạn không cần to chop nội dung vào tiny pieces. Structure, không fragment.
mental models
1. retrieval pipeline — chunk → embed → store → retrieve → generate. Chunking là move một. nếu của bạn nội dung không phải getting cited, hoạt động chain: là nó được crawl, là coherent chunk produced, đã làm đó chunk match query, đã làm nó land nơi LLM sẽ sử dụng nó?
2. Đó chunk là đó unit, không đó trang. Dừng thinking “is my page indexed?” (bản dịch) «là my trang được lập chỉ mục?» và bắt đầu thinking “does my page contain a passage that answers this specific query, on its own?” (bản dịch) «làm my trang contain một passage đó các câu trả lời này cụ thể query, on của nó own?» Lập chỉ mục là necessary; extractability là điều gì wins đó citation.
3. size trade-off — precision so với. context. Nhỏ chunks = sharp nhưng thin. Lớn chunks = rich nhưng noisy. có không universal câu trả lời, và bạn không đặt size anyway — so optimize điều bạn làm control: làm mỗi section coherent đủ để hoạt động tại either granularity.
4. Câu trả lời-đầu tiên beats đó middle. “Lost in the Middle” (bản dịch) «Lost trong đó Middle» says position bên trong đó context window matters. Lead mỗi section với đó definition hoặc claim. Đó giống nhau discipline đó helps human skimmers giữ của bạn point out of đó dead zone.
5. Structure, không fragment. Google nói không chop nội dung vào tiny pieces — so move không phải artificial chunking. nó clear heading hierarchy, một topic per section, self-contained paragraphs. “Chunk optimization” (bản dịch) «Chunk optimization» là mostly chỉ good structure wearing new name.
6. Passage xếp hạng ≠ RAG chunking. giống nhau principle (sub-document granularity), khác outcome. Passage xếp hạng là tín hiệu đó ranks trang; RAG chunking retrieves chunk vào câu trả lời. không conflate SEO-era concept với AI-era một.
Chunk-ready nội dung checklist
truyền để làm của bạn nội dung survive là split, retrieved, và cited out của context:
- mỗi H2/H3 section covers một topic hoặc câu hỏi — không hai-trong-một sections.
- mỗi section leads với câu trả lời (definition / mấu chốt claim đầu tiên, hỗ trợ dưới) — không buried trong paragraph 3–4.
- mỗi major section đọc as self-contained: nó sẽ làm sense trong isolation, với không có gì khoảng nó.
- Sections là appropriate length (~200–500 words) — hoàn tất, không artificially chopped vào tiny blocks.
- phần lớn quan trọng nội dung xuất hiện sớm on trang (Chrome reportedly considers chỉ đầu tiên ~30 passages).
- Heading hierarchy là sạch (logical H1 → H2 → H3) — nó part của structural tín hiệu chunker walks.
- Facts đó phải stay together không phải stranded across có khả năng boundary (e.g. claim trong một paragraph, của nó evidence three paragraphs sau đó).
- được sử dụng các bảng / lists nơi nội dung là genuinely list-like (rõ ràng boundaries chunkers detect; cao hơn citation rates).
- bạn đã làm không rewrite mọi thứ vào rigid word-count blocks — Google nói đó không bắt buộc.
Writing một giant section cho several intents
dài block về definitions, implementation, exceptions, và việc đo lường có thể là hữu ích as trang nhưng noisy as retrieved passage. Split distinct các câu hỏi under headings và cho mỗi section đủ local context để stand alone.
Fragmenting mỗi sentence vào của nó own heading
Tiny chunks có thể lose qualifications và relationships. không optimize cho imagined token number. giữ hoàn tất idea, của nó constraints, và supporting evidence together.
Opening với contextless back-reference
Passages đó begin với “này,” “điều này,” hoặc “tuy nhiên” có thể là separated từ text đó names subject. Open quan trọng sections với trực tiếp sentence đó identifies topic và câu trả lời.
Duplicating text để force overlap
Repeated paragraphs tạo competing near-duplicate passages và tệ hơn reading experience. Retrieval các hệ thống có thể thêm overlap internally; authors nên sử dụng clear transitions và self-contained sections thay vì copying prose.
Prompt: audit passage independence
Review the article section by section as if each section could be retrieved without its
neighbors. For each heading, state the question it answers, whether the opening sentence
names the subject, what context is missing, whether unrelated intents are mixed, and the
smallest edit that makes the section self-contained. Preserve necessary qualifications
and evidence. Do not target an arbitrary token count or rewrite the author's voice.
Article:
[PASTE ARTICLE WITH HEADINGS]Prompt: split overloaded section
This section covers several ideas. Propose a minimal heading structure that groups one
complete intent per section. For each proposed section, write only an answer-first
opening sentence and list which existing paragraphs belong under it. Do not add facts,
remove caveats, duplicate prose, or turn every sentence into a heading.
Section:
[PASTE HEADING AND CONTENT] DevTools Console: flag dài được kết xuất sections
Chạy điều này on bài viết trang. character threshold là review aid, không tìm kiếm-engine chunk boundary.
console.table([...document.querySelectorAll('main h2, main h3')].map((heading, i, all) => {
let text = '';
for (let node = heading.nextElementSibling; node && !all.includes(node); node = node.nextElementSibling) text += ` ${node.textContent}`;
return { heading: heading.textContent.trim(), characters: text.trim().length };
}).filter(row => row.characters > 2000));Regex: tìm contextless section openers trong Markdown
điều này multiline pattern flags headings whose đầu tiên prose word là phổ biến back-reference. Review mỗi hit manually.
^#{2,4}\s+.+\n+(?:\n|>.*\n|\s*)*(This|That|It|They|These|Those|However|Therefore|Also|And|But|So|Then)\b Tools cho passage QA
- trình duyệt outline hoặc document map quickly reveals headings đó combine unrelated các câu hỏi hoặc leave dài stretches không có subheadings.
- vector database hoặc embedding playground có thể demonstrate Cách chunk size ảnh hưởng controlled retrieval kiểm thử, nhưng không convert một model kết quả vào universal SEO prescription.
- Search Console và citation theo dõi đo lường trang outcomes. họ không thể tell bạn chính xác proprietary chunk tìm kiếm nền tảng stored.
Validate chunking-oriented rewrite
| Kiểm thử để chạy | Dự kiến kết quả | thất bại interpretation | Monitoring window | Rollback trigger |
|---|---|---|---|---|
| đọc mỗi edited section không có của nó neighbors | heading và opening identify topic và câu trả lời | passage phụ thuộc vào missing context | Editorial review | Restore context nếu qualification hoặc subject là lost |
| So sánh claims và citations trước khi và sau khi | Facts, caveats, và nguồn relationships vẫn intact | Structural editing changed meaning | trước khi publish | Roll back bất kỳ unsupported hoặc broadened claim |
| Chạy Chunk Tester on old và new versions | Targeted dài hoặc dangling sections improve không có artificial fragmentation | rewrite optimized score thay vì comprehension | trước khi publish | Roll back nếu reading flow hoặc tính đầy đủ worsens |
| Kiểm thử nhỏ versioned retrieval đặt | Relevant sections là retrieved cho dự kiến các câu hỏi không có losing exceptions | Splits đã làm passages cũng thin hoặc mixed intents vẫn | sau khi publish trong controlled hệ thống | Recombine hoặc resplit nếu mấu chốt context repeatedly disappears |
| Inspect được kết xuất heading hierarchy | Headings là ordered, descriptive, và followed by nội dung | Markup thay đổi broke document structure | Phát hành QA | Roll back nếu headings become inaccessible hoặc malformed |
Tự kiểm tra: Chunking
các tài nguyên worth của bạn time
My related writing
- Điều gì chúng ta thực ra Know về Optimizing cho LLM Tìm kiếm — covers Chrome DocumentChunker / 30-passage finding và Cách AI retrieval thực sự xử lý của bạn nội dung.
** foundational research**
- Dense Passage Retrieval (Karpukhin et al., 2020) — Vì sao dense passage retrieval beats từ khóa matching; architecture under modern RAG.
- Retrieval-Augmented Generation (Lewis et al., 2020) — paper đó named RAG; được sử dụng 100-word Wikipedia passages.
- Lost trong Middle (Liu et al., 2023) — LLMs sử dụng bắt đầu và end của context best; Vì sao câu trả lời-đầu tiên matters.
- RAPTOR (Sarthi et al., 2024) — hierarchical/recursive chunking qua tree của summaries.
- là Semantic Chunking Worth Cost? (Vectara, 2024) — study finding semantic chunking không phải reliably tốt hơn hơn fixed-size.
** SEO counterpoint (đọc điều này)**
- SEO Chunk Optimization là Overrated (Despina Gavoyannis, Ahrefs) — case đó “chunk optimization” (bản dịch) «chunk optimization» là mostly chỉ good nội dung structure, và đó Bạn có thể’t control Cách các hệ thống chunk bạn. phần lớn quan trọng nuance on điều này topic.
Practitioner các hướng dẫn
- Pinecone — Chunking Strategies — canonical strategy taxonomy.
- LlamaIndex — Evaluating Ideal Chunk Size (Ravi Theja) — 128/256/512/1024/2048 kiểm thử đó landed on 1 024.
- Databricks — Chunking Strategies cho RAG — six strategies với domain-cụ thể hướng dẫn.
Từ khoảng đó ngành
- Nội dung Chunking Hướng dẫn (Search Engine Land) — covers definition, UX origins, macro/micro/atomic chunk types, và cách chunking connects to AI retrieval.
- Chunk, Cite, Clarify, Xây dựng (Benu Aggarwal, Search Engine Land) — four-part nội dung framework cho AI tìm kiếm; frames “content now competes in a probability-weighted lottery of answer generation.” (bản dịch) «nội dung hiện tại competes trong một probability-weighted lottery of câu trả lời generation.»
- Nội dung Chunking: Điều gì Là Điều này & Nên Bạn Care? (Semrush) — practitioner overview với Mike King quote (“chunking and writing for users is not mutually exclusive” (bản dịch) «chunking và writing cho người dùng không phải mutually exclusive») và Q&MỘT format kiểm thử.
- Chunked, Retrieved, Synthesized (Duane Forrester) — ex-Bing perspective on vì sao structure wins even với lớn context windows; “If traditional SEO optimized for clicks, GenAI systems optimize for chunks.” (bản dịch) «Nếu truyền thống SEO optimized cho clicks, GenAI các hệ thống optimize cho chunks.»
- LLM-Friendly Nội dung (Onely / Bartosz Góralewicz) — chính nguồn cho đó các bảng-increase-citation-rates-2,5x finding và đó listicles-account-cho-50%-of-top-AI-citations stat.
- 2025 AI Citation & LLM Visibility Báo cáo (Đó Digital Bloom) — dữ liệu on điều gì predicts LLM citations; brand tìm kiếm volume, statistics, và quotations all shown to lift visibility.
- Đó Ultimate Hướng dẫn cho Chunking Strategies (Agenta.ai) — comprehensive nhà phát triển-oriented strategy taxonomy với Chroma benchmark dữ liệu across chunking các phương thức.
Stats worth citing
- 9–19% absolute — Cách nhiều dense passage retrieval (DPR) beat BM25 từ khóa matching on top-20 passage độ chính xác. case cho retrieving by meaning, không từ khóa. Karpukhin et al., 2020
- ~7% của queries — share của tìm kiếm queries Google passage xếp hạng ảnh hưởng tại đầy đủ rollout (U.S. English trực tiếp February 10, 2021). Coverage
- +20% absolute độ chính xác — RAPTOR hierarchical-chunking gain on hard QA benchmark Khi paired với GPT-4. Sarthi et al., 2024
- Embedding model > chunking strategy — Vectara 2024 finding đó model quality affected retrieval nhiều hơn hơn liệu chunking là semantic hoặc fixed-size; on thực documents differences là “minimal.” Study
- đầu tiên ~30 passages — number Chrome DocumentChunker reportedly considers per trang (~200 words mỗi), per Dan Petrovic research — argument cho front-loading của bạn mấu chốt nội dung. Qua Ahrefs
- ~2,5x citation rate — Onely finding đó các bảng increase AI-citation rates, với listicles accounting cho ~50% của top AI citations. Structure helps. Onely
- 93,67% của Google AI Overviews cite ít nhất một top-10 organic kết quả — mạnh nhưng không absolute correlation giữa organic thứ hạng và AI citation, meaning passage-level match có thể override xếp hạng position. Digital Bloom, 2025 AI Citation Báo cáo
Nhật ký thay đổi
Đã cập nhật 8 thg 8, 2026.
Tóm tắt biên tập và chi tiết thay đổi đã ghi nhận.Chi tiết thay đổi
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
Không thể so sánh đầy đủ — không có bản lưu trước đó cho lần sửa đổi này.
Đã cập nhật 18 thg 7, 2026.
Tóm tắt biên tập và chi tiết thay đổi đã ghi nhận.Chi tiết thay đổi
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
Không thể so sánh đầy đủ — không có bản lưu trước đó cho lần sửa đổi này.