Hướng dẫn về Chunking

Cách AI các hệ thống split của bạn các trang vào passages to embed, chỉ mục, và retrieve — chunk size, overlap, semantic so với. fixed-size, và cách điều này ties to Google passage xếp hạng.

Xuất bản lần đầu: 24 thg 6, 2026 · Cập nhật lần cuối: 8 thg 8, 2026 · Advanced
Ngôn ngữ

Chunking là đó preprocessing step nơi AI các hệ thống split một document vào nhỏ hơn passages trước embedding và retrieving them. Đó chunk — không đó trang — là đó unit of retrieval trong AI tìm kiếm, so đang được lập chỉ mục không đủ; bạn cần một passage đó các câu trả lời một cụ thể query trong isolation. Chunk size là một trade-off (nhỏ = precise nhưng thin on context; lớn = context-rich nhưng noisy) với không size đó bảo đảm một citation, overlap ngăn boundary mất mát nhưng duplicates tokens, contextual prefixes có thể disambiguate một chunk nhưng mislead retrieval nếu họ là stale, và 'Lost trong đó Middle' — một position effect measured on cụ thể 2023 models và tasks, không mỗi model — có nghĩa là LLMs thường dùng đó bắt đầu và end of context best. Bạn không thể optimize cho một cụ thể chunk size — Google explicitly says bạn không cần to chop nội dung vào pieces — nhưng câu trả lời-đầu tiên, self-contained sections under một clear heading hierarchy chunk và cite cleanly. đây là đó giống nhau sub-document principle as Google passage xếp hạng, applied to RAG.

TL;DR — Chunking là đó preprocessing step đó splits một document vào passages trước họ là embedded, được lập chỉ mục, và retrieved. Điều này tồn tại vì embedding models và context windows có token limits, và vì passage-level retrieval là hơn precise hơn trang-level. Đó chunk, không đó trang, là đó unit of retrieval. Chunk size là một trade-off (nhỏ = precise/thin; lớn = rich/noisy) và không size, overlap giá trị, hoặc splitter bảo đảm một citation; overlap guards boundaries nhưng costs chỉ mục size và duplication; semantic chunking không reliably tốt hơn hơn fixed-size; contextual prefixes có thể disambiguate một chunk nhưng mislead retrieval khi stale. “Lost in the Middle” (bản dịch) «Lost trong đó Middle» — một position effect được ghi lại on cụ thể 2023 models và tasks — có nghĩa là an LLM thường dùng đó bắt đầu và end of của nó context best. Bạn không thể optimize cho một cụ thể chunk size — và Google says bạn không nên try — nhưng câu trả lời-đầu tiên, self-contained sections under một clear heading hierarchy chunk và cite cleanly. đây là đó giống nhau sub-document principle as Google passage xếp hạng, applied to RAG.

Điều gì chunking là, và Vì sao nó tồn tại

Retrieval các hệ thống operate over được lập chỉ mục units, nhưng công khai các công cụ tìm kiếm không expose universal publisher-controlled chunk size. Evidence for this claim Retrieval systems can split files into chunks that are embedded and indexed for later search. Scope: OpenAI's retrieval implementation; chunk sizes, overlap, and indexing behavior are implementation-dependent. Confidence: high · Verified: OpenAI: Retrieval guide RAG research hỗ trợ retrieve-sau đó-generate patterns không có proving đó mỗi AI tìm kiếm sản phẩm dùng giống hệt mechanics. Evidence for this claim Retrieval-augmented generation combines a generator with retrieved external passages or documents. Scope: The original RAG research architecture; it does not establish one optimal chunking strategy for every production system. Confidence: high · Verified: Lewis et al.: Retrieval-Augmented Generation

Chunking là xử lý của dividing document vào nhỏ hơn, discrete segments — chunks hoặc passages — trước khi những điều đó segments là turned vào embeddings, stored trong vector chỉ mục, và retrieved để câu trả lời queries. nó foundational step trong RAG (Retrieval-Augmented Generation) pipelines, và nó part của AI tìm kiếm SEOs underrate phần lớn, vì nó invisible: nó happens during ingestion, không tại query time.

Hai các vấn đề force nó:

  • Đó token-limit vấn đề. Embedding models accept một bounded number of tokens per input — Microsoft notes đó text-embedding-3-small model caps tại 8 191 tokens, và other models là far nhỏ hơn. LLMs có finite context windows cũng. MỘT dài trang đơn giản sẽ không fit as một single unit, so điều này nhận split.
  • Đó retrieval-precision vấn đề. MỘT 5 000-word trang về “technical SEO” (bản dịch) «SEO kỹ thuật» collapsed vào một vector là một fuzzy, averaged tín hiệu. MỘT 400-word section đó là cụ thể về ngân sách crawl, embedded as của nó own vector, là một sharp một. Sub-document granularity là điều gì lets một hệ thống pull đó một relevant passage out of một dài, multi-topic trang.

consequence là single phần lớn quan trọng idea ở đây: ** chunk, không trang, là unit của retrieval trong AI tìm kiếm.** là được crawl và được lập chỉ mục là necessary nhưng chư đủ. bạn cần chunk đó các câu trả lời cụ thể query, rõ ràng, on của nó own. trang xếp hạng #15 organically có thể earn AI citation trong khi #1 kết quả nhận skipped — nếu #15 trang có nhiều hơn extractable passage.

Cách chunking hoạt động, end để end

Chunking changes the retrieval unit: the page is published once, but its passages are stored and matched separately. Nguồn: /ai-search/how-search-works/chunking/

A four-stage flow begins with one long document containing several topics. The system splits and embeds focused passages as separate vectors. A query retrieves one or more best-matching passages, and those selected passages enter the model context for answer generation.

© Patrick Stox LLC · CC BY 4.0 ·

  1. Ingestion + splitting. AI crawler downloads trang; chunking algorithm divides text vào (thường overlapping) segments.
  2. Embedding. mỗi chunk becomes numerical vector và là stored trong vector chỉ mục. (đó embeddings step.)
  3. Retrieval + generation. query là embedded, nearest chunk vectors là pulled qua vector tìm kiếm, và retrieved chunks là đã truyền để LLM đó ghi câu trả lời với citations.

điều này toàn bộ architecture traces back để Dense Passage Retrieval (Karpukhin et al., 2020), mà showed đó retrieving by dense vector similarity beat old từ khóa approach (BM25) by 9–19% absolute on top-20 passage độ chính xác — và đó passage-level retrieval hoạt động tốt hơn hơn document-level cho answering cụ thể các câu hỏi. RAG paper (Lewis et al., 2020) named pattern và được sử dụng 100-word passages từ Wikipedia as của nó chunks. mỗi AI tìm kiếm hệ thống retrieving web nội dung là đang chạy variant của đó pipeline.

Chunking strategies

có không single algorithm — các hệ thống pick từ menu, và họ thay đổi nó theo thời gian:

  • Fixed-size — split by một token hoặc character count, với some overlap. Hầu hết phổ biến. Microsoft ví dụ: “a fixed size sufficient for semantically meaningful paragraphs (for example, 200 words or 600 characters)” (bản dịch) «một fixed size sufficient cho semantically có ý nghĩa paragraphs (ví dụ, 200 words hoặc 600 characters)» với 10–15% overlap.
  • Sentence / paragraph — split on natural language boundaries thay vì an arbitrary token count, bảo toàn semantic units.
  • Semantic / nội dung-aware — group sentences by embedding similarity và split nơi đó topic shifts, aiming to giữ mỗi chunk về một điều.
  • Hierarchical / recursive (RAPTOR) — xây dựng một tree of summaries (document → section → paragraph) so một query có thể là answered tại đó right level of abstraction. RAPTOR (Sarthi et al., 2024) reported một 20% absolute độ chính xác gain on một hard QA benchmark khi paired với GPT-4.
  • Sliding window với overlap — mỗi chunk shares some tokens với của nó neighbors so một sentence split across một boundary không lost.
  • Adaptive / query-dependent (Mix-of-Granularity) — một trained router picks đó chunk size per query. Đó hầu hết sophisticated approach; chưa tiêu chuẩn trong commercial các hệ thống.

Chunk size và overlap — core trade-off

Đây là lever mọi người asks về, và honest câu trả lời là nó phụ thuộc:

  • Nhỏ chunks (128–256 tokens): nhiều hơn precise retrieval, nhưng họ có thể lose xung quanh context câu trả lời cần.
  • Lớn chunks (512–1 024 tokens): nhiều hơn context được bảo toàn, nhưng noisier retrieval — bạn drag trong irrelevant material với relevant bit.

Đó research không crown một winner. LlamaIndex evaluation được tìm thấy 1 024 tokens optimal trong của họ setup; Chroma benchmarks được tìm thấy một 200-token recursive splitter performed consistently across các chỉ số. Ravi Theja takeaway là đó right mindset: “Identifying the best chunk size for a RAG system is as much about intuition as it is empirical evidence.” (bản dịch) «Identifying đó best chunk size cho một RAG hệ thống là as nhiều về intuition as điều này là empirical evidence.» Những numbers mô tả điều gì won trong một team benchmark, on một document set, với một embedding model và evaluation task — không một universal setting. Không chunk size, overlap giá trị, hoặc splitter bảo đảm retrieval, citation, xếp hạng, hoặc inclusion trong an câu trả lời từ an external AI hệ thống; mỗi provider pipeline picks của nó own defaults và có thể thay đổi them không có notice.

Overlap là đó underrated half of này. Không có điều này, nội dung near một boundary nhận split và lost. Microsoft khuyến nghị starting tại 25% overlap so có “smoother transitions between chunks without excessive duplication” (bản dịch) «smoother transitions giữa chunks không có excessive duplication»; other sources suggest 10–15%. Overlap không free, though: overlapping tokens nhận embedded và stored twice, mà inflates đó chỉ mục và có thể surface near-duplicate passages side by side trong một retrieved set — so đây là một trade so với chỉ mục size và redundancy, không một costless safety net. Weigh điều này so với cách thường một boundary thực ra costs bạn một missed fact, không on principle. Đó practical point cho nội dung: không assume đó hệ thống sẽ giữ một key fact intact nếu bạn bury điều này chính xác nơi một chunk có khả năng break.

So, làm semantic chunking luôn win? Không — đó là đó inconvenient finding. Vectara 2024 study được tìm thấy “performance differences are minimal” (bản dịch) «performance differences là minimal» on thực tế documents, và — crucially — đó embedding model quality mattered hơn đó chunking strategy. Khi GPT-4o generated đó các câu trả lời, đó differences giữa strategies đã là “negligible.” Đó kết quả là cụ thể to Vectara document set, models, và evaluation phương thức — đây là evidence semantic chunking không một reliable default win, không proof điều này không bao giờ helps trong any pipeline. Translation cho SEOs: obsessing over an chính xác structure matters far ít hơn hơn nội dung quality và semantic clarity.

Chunk context: Điều gì prefixes có thể (và có thể’t) khắc phục

chunk đó đọc rõ ràng để human có thể vẫn retrieve badly sau khi nó separated từ document khoảng nó — section nó sits under, entity nó thực ra về, qualifier stated hai headings up. Một được ghi lại mitigation là prepending ngắn, chunk-cụ thể context string ( document tiêu đề, section nó belongs để, Điều gì chunk là thực ra về) trước khi chunk là embedded — approach Anthropic calls contextual retrieval. Các tiêu đề, section ancestry, và contextual prefix có thể disambiguate nếu không-orphaned chunk và cut down on retrieval misses gây ra by missing context.

đó khắc phục có thất bại chế độ của của nó own, though: stale hoặc sai context không chỉ fail để help — nó actively misleads retriever toward sai chunk. prefix generated từ heading đó không lâu hơn matches section sau khi trang reorg, hoặc contextual summary đó misstates Điều gì chunk covers, là tệ hơn hơn không prefix tại all. Context preservation là pipeline design lựa chọn với của nó own lỗi chế độ, không một-way improvement Bạn có thể bolt on và forget.

Google passage xếp hạng — SEO ancestor của chunking

Google đã được đang làm sub-document granularity since dài trước “RAG” đã là một buzzword. Tại Tìm kiếm On 2020, Prabhakar Raghavan announced passage xếp hạng: “By better understanding the relevancy of specific passages, not just the overall page, we can find that needle-in-a-haystack information you’re looking for.” (bản dịch) «By tốt hơn understanding đó relevancy of cụ thể passages, không chỉ đó overall trang, we có thể tìm đó needle-trong-một-haystack information bạn là looking cho.» Điều này went trực tiếp trong U.S. English on February 10, 2021 và ảnh hưởng roughly 7% of queries.

Hai điều mọi người nhận sai về nó:

  • đây là “passage ranking,” (bản dịch) «passage xếp hạng,» không “passage indexing.” (bản dịch) «passage lập chỉ mục.» Google đầu tiên announcement dùng “lập chỉ mục,” thì quickly corrected điều này: “this change doesn’t mean we’re indexing individual passages independently of pages.” (bản dịch) «này thay đổi không có nghĩa là chúng ta là lập chỉ mục individual passages independently of các trang.» Đó trang là vẫn được lập chỉ mục as một toàn bộ; đó relevant passage là an additional tín hiệu xếp hạng.
  • Đó trang ranks, không đó passage. John Mueller: “Passage ranking is not about ranking a specific passage but understanding the content on a really long, not SEO optimized page, and ranking that page (not the passage) for a query where the passage is relevant.” (bản dịch) «Passage xếp hạng không phải về xếp hạng một cụ thể passage nhưng understanding đó nội dung on một thực sự dài, không SEO optimized trang, và xếp hạng đó trang (không đó passage) cho một query nơi đó passage là relevant.»

So passage xếp hạng và RAG chunking share principle — paragraph, không trang, là thường right unit để match so với cụ thể query — nhưng outcome differs: passage xếp hạng lifts trang xếp hạng; RAG chunking retrieves chunk để feed generated câu trả lời. giống nhau idea, khác machinery. ( kỹ thuật underpinning, per Dawn Anderson reporting, là DeepCT — BERT-derived contextual term weights replacing TF-IDF, so term frequency không lâu hơn equals term relevance.)

”Lost in the Middle” (bản dịch) «Lost trong đó Middle» — nơi của bạn câu trả lời sits matters

Even sau của bạn chunk là retrieved, nơi điều này lands trong đó LLM context ảnh hưởng liệu đó model thực ra dùng điều này. Đó Stanford “Lost in the Middle” (bản dịch) «Lost trong đó Middle» study (Liu et al., 2023) được tìm thấy đó “performance is often highest when relevant information occurs at the beginning or end of the input context, and significantly degrades when models must access relevant information in the middle of long contexts.” (bản dịch) «performance là thường highest khi relevant information occurs tại đó beginning hoặc end of đó input context, và significantly degrades khi models phải access relevant information trong đó middle of dài contexts.»

đó finding không apply universally theo mặc định — nó position effect measured on named multi-document QA và mấu chốt-giá trị retrieval tasks, on cụ thể model generation Liu et al. tested trong 2023. nó không proof đó mỗi hiện tại model bỏ qua evidence placed centrally trong của nó context; khác architectures, lâu hơn effective context windows, và newer training có thể hẹp hoặc widen effect. Treat nó as được ghi lại risk để design khoảng, không fixed law của mỗi model bạn’ll bao giờ là retrieved vào.

nội dung implication là vẫn concrete và thấp-risk either way: lead với của bạn câu trả lời. Put definition, Điểm mấu chốt finding, trực tiếp câu trả lời trong đầu tiên sentence của mỗi section — không buried trong paragraph four. Đây là giống nhau câu trả lời-đầu tiên (BLUF) discipline đó phục vụ human skimmers; nó chỉ happens để cũng hedge so với là dropped vào middle của context window on models nơi effect holds.

practical constraints Bạn có thể’t see

  • Chrome ~30-passage limit. Dan Petrovic research suggests Chrome DocumentChunker analyzes nội dung trong ~200-word passages và “only ever considers the first 30 passages of a page.” (bản dịch) «chỉ bao giờ considers đó đầu tiên 30 passages of một trang.» Điều này tree-walks đó semantic HTML top to bottom. Implication: của bạn hầu hết quan trọng nội dung nên xuất hiện sớm — không tại đó bottom of một 10 000-word trang.
  • Bạn không control đó chunker. Despina Gavoyannis (Ahrefs) là blunt: “You can’t control how Google, ChatGPT, or Perplexity chunk your content. Their pipelines change based on cost, model, and context.” (bản dịch) «Bạn không thể control cách Google, ChatGPT, hoặc Perplexity chunk nội dung của bạn. Của họ pipelines thay đổi dựa trên cost, model, và context.» Và: “Manual ‘chunk optimization’ is impossible in practice.” (bản dịch) «Manual ‘chunk optimization’ là không thể trong thực tế.»

Điều gì “chunk optimization” (bản dịch) «chunk optimization» thực ra là (và Google caveat)

Ở đây đó nuance đó hype skips. Google 2026 AI optimization hướng dẫn says plainly: “There’s no requirement to break your content into tiny pieces for AI to better understand it.” (bản dịch) «có không requirement to break nội dung của bạn vào tiny pieces cho AI to tốt hơn understand điều này.» Của nó các hệ thống “are able to understand the nuance of multiple topics on a page and show the relevant piece to users.” (bản dịch) «là able to understand đó nuance of multiple topics on một trang và cho thấy đó relevant piece to người dùng.» So không — làm không rewrite nội dung của bạn vào rigid 300-word blocks.

Nhưng đó không có nghĩa là structure là irrelevant. As Gavoyannis puts điều này, “Most SEOs using the term [chunk optimization] are just talking about good content structure.” (bản dịch) «Hầu hết SEOs dùng đó term [chunk optimization] là chỉ talking về good nội dung structure.» Đó advice underneath đó buzzword là sound; đó framing chỉ inflates của nó novelty. Duane Forrester line captures đó shift well: “If traditional SEO optimized for clicks, GenAI systems optimize for chunks… Structure still wins.” (bản dịch) «Nếu truyền thống SEO optimized cho clicks, GenAI các hệ thống optimize cho chunks… Structure vẫn wins.»

So Điều gì làm bạn thực ra làm? Ghi chunk-ready nội dung:

  • Một topic per section. focused H2/H3 maps cleanly để coherent chunk.
  • Câu trả lời-đầu tiên. Lead với claim; hỗ trợ nó dưới.
  • Self-contained sections. Kiểm thử: sẽ điều này paragraph làm sense nếu nó appeared trong isolation? nếu không, nó sẽ không survive là chunked out.
  • Appropriate length. 200–500 words per major section lines up naturally với 256–512-token chunks — hoàn tất đủ để là hữu ích, không artificially ngắn.
  • Structured formats. Các bảng và lists cho các hệ thống rõ ràng boundaries để detect. (Onely research: các bảng increase citation rates ~2,5x.)

Mike King framing là đó right reassurance: “chunking and writing for users is not mutually exclusive.” (bản dịch) «chunking và writing cho người dùng không phải mutually exclusive.» Đó structure đó helps một reader skim là đó structure đó chunks cleanly. bạn là không optimizing cho một robot tại đó expense of một human — đây là đó giống nhau nội dung.

Passage xếp hạng so với. RAG chunking — side by side

Google passage xếp hạngRAG chunking
Điều gì nó làxếp hạng tín hiệupreprocessing step
GranularityPassage trong trangChunk split trước khi embedding
Outcometrang ranks cao hơnchunk là retrieved vào câu trả lời
nơi nó chạyTại xếp hạng timeTại ingestion (sau đó retrieval)
bạn control split?KhôngKhông
Shared principleSub-document granularity — paragraph, không trang, thường matches cụ thể query best

nơi điều này sits trong pipeline

Chunking là đầu tiên move trong AI retrieval: chunk → embed → store → retrieve → generate. nó feeds embeddings (mỗi chunk becomes vector), mà feed vector tìm kiếm (query matches nearest chunks), mà feeds RAG (retrieved chunks become câu trả lời). Upstream, AI các crawler là Cách của bạn nội dung nhận ingested trong đầu tiên place. cho truyền thống-tìm kiếm version của điều này toàn bộ pipeline, see Cách Tìm kiếm Hoạt động.

Add an expert note

Pin an expert quote

New person? Create their unclaimed profile at /admin/experts/ → Pin a quote first.