Tokens và Context Windows
Điều gì tokens và context windows là, vì sao LLMs có them, và vì sao đó context-window limit là đó mechanical reason AI tìm kiếm chunks và retrieves nội dung của bạn thay vì reading toàn bộ các trang.
Ngôn ngữ
MỘT token là đó smallest unit of text an LLM xử lý — khoảng 4 characters, hoặc về ¾ of an English word (100 tokens ≈ 60–80 words theo Google). Tokenization splits text (và images, audio, video) vào những units. MỘT context window là đó maximum number of tokens một model có thể hold tại khi — input plus output combined — như ngắn-term memory; Google Gemini có thể accept lên để 1 million tokens. Nội dung bên ngoài đó window là invisible, không dần forgotten. Này matters cho AI tìm kiếm vì không đầy đủ trang hoặc site fits trong một window, so retrieval các hệ thống select và chunk nội dung của bạn trước an LLM bao giờ sees điều này — mà là đó mechanical reason chunking, passage-cấp độ retrieval, và 'giữ điều này clear và self-contained' advice exist. MỘT bigger context window không có nghĩa là an AI đọc của bạn toàn bộ site, và Google says có không ideal trang length và không requirement để fragment nội dung cho AI.
TL;DR — MỘT token là đó little piece of text an AI đọc — thường một chunk of một word, khoảng ¾ of an English word. MỘT context window là cách nhiều text an AI có thể hold trong của nó head tại khi (của nó ngắn-term memory). Cả hai quan trọng cho tìm kiếm vì tìm kiếm các hệ thống không routinely hand một model mỗi trang hoặc an entire website. They thường grab chỉ đó relevant parts trước they câu trả lời.
Điều gì một token là
Language models xử lý text as tokens, mà có thể là toàn bộ words hoặc nhỏ hơn character sequences depending on đó tokenizer. Evidence for this claim OpenAI models process text as tokens, and token boundaries may be whole words or parts of words. Scope: OpenAI tokenization; token counts depend on the model-compatible encoding. Confidence: high · Verified: OpenAI: What are tokens? MỘT context window là một model-cụ thể limit on đó tokens khả dụng để một yêu cầu và phản hồi, và published limits có thể thay đổi by model version. Evidence for this claim Context-window limits are documented per model and can differ across model versions. Scope: OpenAI model metadata; published limits are product-specific and can change. Confidence: high · Verified: OpenAI: Models
Lớn language models — đó tech behind ChatGPT, Gemini, và AI Overviews — không đọc text đó way bạn làm, word by word. They break điều này vào tokens: nhỏ units đó là thường part of một word thay vì một toàn bộ một. Google rule of thumb cho của nó Gemini models là đó một token là về 4 characters, và 100 tokens là khoảng 60–80 English words. So một token là một bit ít hơn một đầy đủ word on average.
Đó xử lý of chopping text vào tokens là called tokenization, và đó đầy đủ set of tokens một model knows là của nó vocabulary. Phổ biến words có thể là một single token; lâu hơn hoặc unusual words nhận split vào several. Và đây là không chỉ text — Gemini tokenizes images, audio, và video cũng.
Điều gì một context window là
MỘT context window là đó hầu hết tokens một model có thể hoạt động với trong một go. Think of điều này as ngắn-term memory — đó là Google own analogy. Google mô tả đó space as một shared pool covering đó instructions đó hệ thống cho đó model, đó lại-và-forth of của bạn conversation, bất kỳ documents đó đã nhận pulled trong, của bạn câu hỏi, và đó model câu trả lời — though đó chính xác split không giống hệt mọi nơi; some providers cap cách nhiều output bạn có thể nhận lại riêng từ đó overall window, so kiểm tra đó specifics cho whatever model bạn là thực ra dùng.
Phần quan trọng: khi điều gì đó falls bên ngoài đó usable window, đây là đã biến mất — không slowly forgotten, chỉ invisible, as nếu điều này đã là không bao giờ ở đó. Đó model không thể reach lại cho điều này trừ khi đó hệ thống có chủ ý re-feeds điều này, hoặc đó sản phẩm thay vì rejects, truncates, hoặc compacts của bạn yêu cầu trước đó model ngay cả sees điều này.
Vì sao này matters cho tìm kiếm
Ở đây đó connection mọi người miss. Ngay cả khi một trang có thể fit bên trong một lớn context window, handing over mỗi trang hoặc an entire site là expensive và thường irrelevant để đó câu hỏi. So khi an AI tìm kiếm tool các câu trả lời, điều này generally selects và chunks đó hầu hết relevant passages đầu tiên, thì feeds chỉ những để đó model.
đó là đó toàn bộ reason bạn hear advice như “put your answer up front” (bản dịch) «put của bạn câu trả lời lên front» và “keep sections self-contained.” (bản dịch) «giữ sections self-contained.» đây là không một style preference — đây là vì đó AI là hoạt động với một piece of trang của bạn, không đó trang. Nếu một section chỉ làm hợp lý trong đó context of đó rest of đó trang, điều này có thể lose của nó meaning đó moment đây là pulled out on của nó own.
Đó điều mọi người nhận sai
MỘT giant context window không có nghĩa là an AI đọc của bạn toàn bộ site. Bạn’ll hear đó Gemini có thể take trong một million tokens, và think “great, it’ll read everything I publish.” (bản dịch) «great, điều này’ll đọc mọi thứ I publish.» Điều này sẽ không. Trong một tìm kiếm hoặc AI Overview scenario, đó hệ thống vẫn picks và chunks relevant passages trước bất cứ điều gì reaches đó model — đó huge window thay đổi điều gì là có thể trong principle, không điều gì một công cụ tìm kiếm thực ra feeds itself cho của bạn query.
Và bạn không cần để pre-cut nội dung của bạn vào tiny token-sized chặn. Google says outright có không ideal trang length và không requirement để break nội dung vào tiny pieces cho AI. Ghi cho mọi người; đó chunking happens on đó machine side.
Muốn đó real mechanics — cách tokenization hoạt động, cách big context windows có gotten và vì sao có một ceiling, và điều gì “Lost in the Middle” (bản dịch) «Lost trong đó Middle» có nghĩa là cho của bạn nội dung? Chuyển để đó Advanced tab.
TL;DR — MỘT token là đó unit an LLM xử lý — một sub-word fragment, ~4 characters, với 100 tokens ≈ 60–80 English words (Google). Tokenization splits all input và output — text, images, audio, video — vào tokens; đó model known set là của nó vocabulary. MỘT context window là đó total token budget shared by input (hệ thống prompt + history + retrieved tài liệu + của bạn query) và output (đó phản hồi) — Google analogy là ngắn-term memory. Nội dung bên ngoài đó window là invisible, không dần forgotten. Windows scaled từ ~2K tokens để 1M+ (Gemini), với một hardware ceiling (“thermal limit” (bản dịch) «thermal limit» of TPUs). Đó reason này matters cho tìm kiếm: các hệ thống manage cost và relevance by selecting và chunking trước an LLM sees của bạn nội dung — và “Lost in the Middle” (bản dịch) «Lost trong đó Middle» có nghĩa là ngay cả điều gì là trong đó window không dùng evenly. MỘT bigger window không có nghĩa là AI đọc của bạn toàn bộ site.
Điều gì một token thực ra là
Tokenization maps text vào integer token IDs dùng một model-compatible encoding; token boundaries không phải đó giống nhau as word boundaries. Evidence for this claim OpenAI models process text as tokens, and token boundaries may be whole words or parts of words. Scope: OpenAI tokenization; token counts depend on the model-compatible encoding. Confidence: high · Verified: OpenAI: What are tokens? Context-window sizes là sản phẩm và model metadata, không một vĩnh viễn thuộc tính of all language models. Evidence for this claim Context-window limits are documented per model and can differ across model versions. Scope: OpenAI model metadata; published limits are product-specific and can change. Confidence: high · Verified: OpenAI: Models
Google là blunt về đó granularity: “Gemini and other generative AI models process input and output at a granularity called a token.” (bản dịch) «Gemini và other generative AI models xử lý input và output tại một granularity called một token.» MỘT token không phải một word và không một character — đây là một fragment. As Google tài liệu put điều này, “Long words are broken up into several tokens. The set of all tokens used by the model is called the vocabulary, and the process of splitting text into tokens is called tokenization.” (bản dịch) «Dài words là hỏng lên vào several tokens. Đó set of all tokens dùng by đó model là called đó vocabulary, và đó xử lý of splitting text vào tokens là called tokenization.»
Đó citable rule of thumb, straight từ Google: “For Gemini models, a token is equivalent to about 4 characters. 100 tokens is equal to about 60-80 English words.” (bản dịch) «Cho Gemini models, một token là tương đương để về 4 characters. 100 tokens là equal để về 60-80 English words.» So khoảng ¾ of một word theo token, on average — nhưng average là đang làm hoạt động ở đó. Punctuation, non-English text, code, và unusual hoặc dài words tokenize ít hơn efficiently (hơn tokens theo word), mà là chính xác vì sao counting words là một poor way để estimate token usage.
Tokens cũng không chỉ text. Google: “All input to and output from the Gemini API is tokenized, including text, image files, and other non-text modalities.” (bản dịch) «All input để và output từ đó Gemini API là tokenized, including text, image files, và other non-text modalities.» Images, audio, và video frames all become tokens cũng. Google DeepMind engineers mô tả đó context-window đo lường as “how many tokens — the smallest building blocks, like part of a word, image or video — that the model can process at once.” (bản dịch) «cách nhiều tokens — đó smallest building chặn, như part of một word, image hoặc video — đó model có thể xử lý tại khi.»
Điều gì một context window là
MỘT context window là đó maximum number of tokens một model có thể hold và reason over trong một single interaction. Google cách diễn đạt là có chủ ý đơn giản: “An analogy for the context window is short term memory.” (bản dịch) «An analogy cho đó context window là ngắn term memory.»
Đó key mechanical point là đó đây là một shared budget. Mọi thứ competes cho đó giống nhau space:
- đó hệ thống prompt (đó instructions đó application cho đó model),
- đó conversation history so far,
- bất kỳ retrieved hoặc injected documents (đó passages một tìm kiếm/RAG hệ thống pulled trong),
- của bạn input (đó hiện tại query), và
- đó model own output (đó câu trả lời điều này generates).
Google Gemini tài liệu mô tả input và output drawing từ một shared pool. Đó chính xác accounting không universal, though: other providers publish một tách biệt maximum- output hình alongside đó context window thay vì treating điều này as một undivided budget — Anthropic model tài liệu, chẳng hạn, list một “context window” (bản dịch) «context window» size và một distinct “max output” cap side by side cho mỗi Claude model. Evidence for this claim Context-window limits are documented per model and can differ across model versions. Scope: OpenAI model metadata; published limits are product-specific and can change. Confidence: high · Verified: OpenAI: Models Kiểm tra đó tài liệu cho đó cụ thể model và sản phẩm bạn là thực ra dùng rather hơn assuming một formula áp dụng mọi nơi.
Whatever đó chính xác accounting, khi nội dung falls bên ngoài đó usable window, điều này không dần forgotten — đây là đơn giản invisible, as nếu điều này không bao giờ existed, trừ khi một hệ thống explicitly summarizes điều này và re-injects một ngắn hơn version, hoặc đó sản phẩm thay vì rejects, truncates, hoặc compacts đó yêu cầu trước đó model bao giờ sees điều này. Đó “it’s gone, not fading” (bản dịch) «đây là đã biến mất, không fading» behavior là đó part đó trips mọi người lên khi they assume an AI “remembers” một dài conversation đó way một person sẽ.
A page or corpus contains many possible passages. Retrieval selects the passages most relevant to the current question. Those passages then share a model-specific context budget with system instructions, conversation history, the current query, and an output allowance. A larger window increases capacity, but it does not mean a search system routinely sends a whole page or site to the model.
© Patrick Stox LLC · CC BY 4.0 ·
Cách big context windows là — và cách fast đó changed
Đó quy mô-lên ở đây đã được dramatic. Sớm generative models handled chỉ một couple thousand tokens; đó progression ran qua khoảng 8K, thì 32K, thì 128K, trước đó jump Google flags trong của nó tài liệu: “Gemini is the first model capable of accepting 1 million tokens.” (bản dịch) «Gemini là đó đầu tiên model capable of accepting 1 million tokens.»
có một physical ceiling, though. Google DeepMind pushed để experimental 10-million- token windows, và research scientist Nikolay Savinov — một of đó research leads on đó dài-context project, ai originally targeted 128 000 tokens trước landing on 1 million — put đó limit plainly: “10 million tokens at once is already close to the thermal limit of our Tensor Processing Units.” (bản dịch) «10 million tokens tại khi là đã close để đó thermal limit of của chúng ta Tensor Processing Units.» Bigger context không free — điều này costs memory, compute, và theo nghĩa đen heat. (Google DeepMind Denis Teplyashin có described đó giống nhau cascade từ 128K để 512K để 1M để 10M, mỗi step opening new possibilities nhưng đang chạy vào harder engineering limits; research scientist Machel Reid có described điều gì nhóm thực ra làm với đó space, như feeding an entire codebase hoặc một 45-minute film vào một single prompt.)
Hai consequences worth internalizing:
- Đó numbers bạn see là moving targets. “The largest context window” (bản dịch) «Đó largest context window» là một hình đó giữ thay đổi; không hard-code strategy để một cụ thể token count.
- Bigger ≠ tự động tốt hơn. Hơn capacity không phải hơn comprehension — see “Lost in the Middle” (bản dịch) «Lost trong đó Middle» dưới.
Điều gì một million tokens thực ra looks như
Abstract token được tính là hard để feel, so Google offers concrete sizing cho điều gì 1M tokens có thể hold:
- “50,000 lines of code (with the standard 80 characters per line),” (bản dịch) «50 000 lines of code (với đó tiêu chuẩn 80 characters theo line),»
- “All the text messages you have sent in the last 5 years,” (bản dịch) «All đó text messages bạn có đã gửi trong đó cuối cùng 5 năm,»
- “8 average length English novels,” (bản dịch) «8 average length English novels,»
- “Transcripts of over 200 average length podcast episodes.” (bản dịch) «Transcripts of over 200 average length podcast episodes.»
đó là genuinely enormous — và đây là chính xác vì sao đó “so the AI just reads my whole site” (bản dịch) «so đó AI chỉ đọc my toàn bộ site» assumption feels reasonable và là vẫn sai. Mà brings us để đó tìm kiếm angle.
Vì sao tokens và context windows quan trọng cho AI tìm kiếm
Này là đó cốt lõi of điều này cho anyone đang làm SEO hoặc nội dung. Không đầy đủ trang — và certainly không đầy đủ site — là reliably handed để một model toàn bộ trong một tìm kiếm scenario. Ngay cả khi một trang fits technically, feeding all of điều này có thể là uneconomical hoặc irrelevant để đó query. Đó là vì sao retrieval các hệ thống — AI Overviews, Copilot, RAG pipelines, chatbots với browsing — select, chunk, và truyền chỉ some of một trang nội dung để đó model. (Không tìm kiếm vendor publishes chính xác điều gì điều này retrieves theo query, so này là inference từ cách retrieval architectures hoạt động generally và từ công khai statements như đó ones dưới — không một claim đó mỗi AI tìm kiếm sản phẩm behaves identically.)
Đó token/context-window limit là đó mechanical reason behind một toàn bộ stack of AI- tìm kiếm behavior bạn đã know về on này site:
- đây là vì sao chunking tồn tại — nội dung nhận split vào retrievable passages vì đó toàn bộ điều sẽ không fit, và Microsoft own Azure hướng dẫn says partitioning lớn documents vào nhỏ hơn chunks “can help you stay under the maximum token input limits of chat completion and embedding models.” (bản dịch) «có thể help bạn stay dưới đó maximum token input limits of chat completion và embedding models.»
- đây là vì sao retrieval (đó “R” trong RAG) có để pick một handful of passages trước generation — đó model có thể chỉ là handed điều gì fits trong của nó budget.
- đây là vì sao embedding models có của họ own token caps (nhiều top out khoảng vài thousand tokens theo input), so dài passages có để là split trước họ là ngay cả turned vào vectors.
Đó ngành own cách diễn đạt of đó vấn đề là rõ ràng. Đó llms.txt spec opens với: “Large language models increasingly rely on website information, but face a critical limitation: context windows are too small to handle most websites in their entirety.” (bản dịch) «Lớn language models increasingly rely on website information, nhưng face một cốt yếu limitation: context windows là cũng nhỏ để xử lý hầu hết websites trong của họ entirety.» đó là đó toàn bộ motivation cho đó format — ngay cả khi, as I’ll cover trong đó myths section, Google John Mueller có called llms.txt tại best một token-saving “crutch” cho coding tools, không một tìm kiếm-visibility mechanism.
”Lost in the Middle” (bản dịch) «Lost trong đó Middle» — ngay cả điều gì là trong đó window không dùng evenly
Ở đây đó subtler point. Getting nội dung của bạn vào đó window không đó finish line. Đó Stanford “Lost in the Middle” (bản dịch) «Lost trong đó Middle» nghiên cứu (Liu et al., 2023) được tìm thấy đó on đó language models và tasks they tested, models dùng information tại đó bắt đầu và end of của họ context far tốt hơn information buried trong đó middle. đó là một position effect measured on 2023-era models — không proof mỗi hiện tại model luôn bỏ qua đó middle — nhưng đây là held lên as một chung caution: stuffing hơn tokens trong không bảo đảm tốt hơn các câu trả lời, và nội dung đó leads với của nó point survives cả hai retrieval và trong-context attention tốt hơn nội dung đó buries điều này.
Đó economics: vì sao AI các hệ thống là selective on purpose
có một business-model reason retrieval các hệ thống chunk aggressively thay vì feed toàn bộ các trang. Google notes đó “when billing is enabled, the cost of a call to the Gemini API is determined in part by the number of input and output tokens” (bản dịch) «khi billing là enabled, đó cost of một call để đó Gemini API là determined trong part by đó number of input và output tokens» — và output tokens typically cost several times hơn input tokens. Tokens là theo nghĩa đen metered. Đó cho AI các hệ thống một mạnh economic incentive để retrieve và truyền chỉ điều gì là necessary, mà reinforces đó giống nhau conclusion từ một khác nhau direction: selective, passage-cấp độ retrieval không một tạm thời limitation để chờ out — đây là cách những các hệ thống là designed để hoạt động.
Phổ biến myths, debunked
-
“A 1-million-token context window means the AI reads my whole website at once.” (bản dịch) «MỘT 1-million-token context window có nghĩa là đó AI đọc my toàn bộ website tại khi.» Không. Trong một tìm kiếm hoặc RAG scenario, retrieval vẫn selects và chunks nội dung trước điều này bao giờ reaches đó model window. MỘT huge window thay đổi điều gì là có thể trong principle, không điều gì một công cụ tìm kiếm hoặc AI Overview thực ra feeds itself theo query. Google line là trực tiếp on point: “There’s no ideal page length, and in the end, make pages for your audience, not just for generative AI search.” (bản dịch) «có không ideal trang length, và trong đó end, làm các trang cho của bạn audience, không chỉ cho generative AI tìm kiếm.»
-
“I need to write in exactly 200-word, token-sized chunks.” (bản dịch) «I cần để ghi trong chính xác 200-word, token-sized chunks.» Không. Google: “There’s no requirement to break your content into tiny pieces for AI to better understand it.” (bản dịch) «có không requirement để break nội dung của bạn vào tiny pieces cho AI để tốt hơn understand điều này.» Chunking happens on đó hệ thống side; của bạn job là clear, well- structured nội dung, không manual token accounting.
-
“More tokens = a smarter model / better answers.” (bản dịch) «Hơn tokens = một smarter model / tốt hơn các câu trả lời.» Không. MỘT lớn hơn window increases capacity, không comprehension — và “Lost in the Middle” (bản dịch) «Lost trong đó Middle» cho thấy models dùng đó bắt đầu và end of của họ context tốt hơn đó middle. Bigger không tự động tốt hơn.
-
“Tokens = words, so I can just count words to estimate usage.” (bản dịch) «Tokens = words, so I có thể chỉ count words để estimate usage.» Khoảng, nhưng không reliably. 100 tokens ≈ 60–80 English words là Google average, nhưng punctuation, code, non-English text, và unusual words tokenize ít hơn efficiently — mà matters khi bạn là estimating cost hoặc budget.
-
“llms.txt solves the context-window problem for my site.” (bản dịch) «llms.txt solves đó context-window vấn đề cho my site.» Đó vấn đề là real — đây là đó spec own stated motivation — nhưng Google không dùng llms.txt cho tìm kiếm, và Mueller có được diễn đạt điều này as một token-saving crutch cho AI coding tools, không an SEO cách sửa. Đó mitigations đó thực ra quan trọng là đó giống nhau fundamentals: là crawlable và được lập chỉ mục, và ghi clear, retrieval-friendly nội dung.
Điều gì này có nghĩa là cho nội dung và SEO
Strip đó jargon và đó takeaways là concrete — và họ là đó giống nhau discipline đó cho thấy lên trên này cluster:
- Front-load đó câu trả lời. Cả hai retrieval và trong-context attention favor đó bắt đầu (và end) of điều gì một model sees. Lead với của bạn point.
- Giữ sections self-contained. Vì một passage có thể là pulled out of trang của bạn và handed để một model on của nó own, điều này nên stand on của nó own. Này là đó giống nhau logic behind good H2/H3 structure và clear topic sentences.
- không obsess over manual chunking. Không ideal trang length, không requirement để fragment. Ghi cho humans; đó hệ thống chunks.
- “Bigger context window” (bản dịch) «Bigger context window» ≠ “AI reads my whole site.” (bản dịch) «AI đọc my toàn bộ site.» Optimize cho đang được tìm thấy và được chọn, không cho đang ingested toàn bộ.
Tokens và context windows là đó thấp-cấp độ plumbing dưới hầu hết of này cluster: họ là đó reason grounding và retrieval có để select passages, đó reason chunking prepares text đó way điều này làm, đó constraint embeddings và semantic và vector tìm kiếm hoạt động trong, và đó limit passage xếp hạng tồn tại để hoạt động khoảng. Và they sit right tiếp theo để đó model other hard boundary — của nó knowledge cutoff — as một of đó cốt lõi limitations of bất kỳ LLM.
AI summary
MỘT condensed take on đó Advanced version:
- MỘT token = đó unit an LLM xử lý — một sub-word fragment, ~4 characters, với 100 tokens ≈ 60–80 English words (Google). Tokenization splits all input và output — text, images, audio, video — vào tokens; đó model known set là của nó vocabulary.
- Word được tính estimate tokens poorly — punctuation, code, non-English text, và unusual words tokenize ít hơn efficiently (hơn tokens theo word).
- MỘT context window = đó total token budget shared by input (hệ thống prompt + history + retrieved tài liệu + query) và output. Google analogy: ngắn-term memory.
- Nội dung bên ngoài đó window là invisible, không dần forgotten — đã biến mất trừ khi một hệ thống summarizes và re-injects điều này.
- Windows scaled fast: ~2K → 8K/32K/128K → 1M+ (Gemini), với experimental 10M near đó “thermal limit” (bản dịch) «thermal limit» of Google TPUs (Nikolay Savinov). Đó numbers giữ moving.
- Bigger ≠ tốt hơn: hơn capacity không hơn comprehension. “Lost in the Middle” (bản dịch) «Lost trong đó Middle» — models dùng đó bắt đầu và end of của họ context tốt hơn đó middle.
- Vì sao điều đó quan trọng cho tìm kiếm: không đầy đủ trang hoặc site fits trong một window, so retrieval các hệ thống (AI Overviews, Copilot, RAG) select và chunk trước an LLM bao giờ sees của bạn nội dung. đây là đó mechanical reason chunking, retrieval, và embedding token caps exist.
- Economics reinforce điều này: tokens là metered (output thường costs hơn input), so các hệ thống là designed để retrieve selectively, không feed toàn bộ các trang.
- Myths busted: một 1M window không có nghĩa là AI đọc của bạn toàn bộ site; không cần cho manual token-sized chunks (Google says không ideal length, không requirement để fragment); hơn tokens ≠ smarter; llms.txt không solve điều này cho tìm kiếm.
- Nội dung upshot: front-load đó câu trả lời, giữ sections self-contained, không obsess over manual chunking, optimize cho đang được chọn, không ingested toàn bộ.
Tài liệu chính thức
Chính-nguồn tài liệu on tokens và context windows.
- Understand và count tokens (Gemini API) — điều gì một token là, đó ~4-characters / 100-tokens-≈-60-80-words rule of thumb, tokenization và vocabulary, và multimodal tokenization.
- Dài context (Gemini API) — đó ngắn-term-memory analogy, đó 1-million-token milestone, đó lịch sử progression, và đó concrete “what a million tokens looks like” (bản dịch) «điều gì một million tokens looks như» các ví dụ.
- Điều gì là một dài context window? Google DeepMind engineers giải thích (Google Blog) — DeepMind engineers on tokens as đó smallest building chặn và đó hardware ceiling.
- Gemini Nhà phát triển API pricing — token-based billing; input so với. output token costs.
- Optimizing của bạn website cho generative AI features (Google Search Central) — “no ideal page length” (bản dịch) «không ideal trang length» và “no requirement to break your content into tiny pieces.” (bản dịch) «không requirement để break nội dung của bạn vào tiny pieces.»
Microsoft / Azure
- Chunk lớn documents cho vector tìm kiếm (Azure AI Tìm kiếm) — chunking để stay dưới đó maximum token input limits of chat-completion và embedding models.
- Điều gì là AI Tokens? (Microsoft Copilot help) — Microsoft consumer-facing explainer of tokens và cách token limits cap cách nhiều bạn có thể ask về tại khi.
Ngành vấn đề statement
- llms.txt spec (llmstxt.org) — frames đó context-window limit as đó origin vấn đề: windows “too small to handle most websites in their entirety.” (bản dịch) «cũng nhỏ để xử lý hầu hết websites trong của họ entirety.»
Quotes từ đó nguồn
On-đó-record statements từ Google và Microsoft, plus đó ngành cách diễn đạt of đó context-window vấn đề. Mỗi link là một deep link để đó quoted passage nơi đó nguồn trang hỗ trợ một.
Google — điều gì một token là
- “Gemini and other generative AI models process input and output at a granularity called a token.” (bản dịch) «Gemini và other generative AI models xử lý input và output tại một granularity called một token.» — Google AI cho Nhà phát triển, Understand và count tokens. Nhảy đến trích dẫn
- “For Gemini models, a token is equivalent to about 4 characters. 100 tokens is equal to about 60-80 English words.” (bản dịch) «Cho Gemini models, một token là tương đương để về 4 characters. 100 tokens là equal để về 60-80 English words.» Nhảy đến trích dẫn
- “Long words are broken up into several tokens. The set of all tokens used by the model is called the vocabulary, and the process of splitting text into tokens is called tokenization.” (bản dịch) «Dài words là hỏng lên vào several tokens. Đó set of all tokens dùng by đó model là called đó vocabulary, và đó xử lý of splitting text vào tokens là called tokenization.» Nhảy đến trích dẫn
Google — đó context window
- “An analogy for the context window is short term memory.” (bản dịch) «An analogy cho đó context window là ngắn term memory.» — Google AI cho Nhà phát triển, Dài context. Nhảy đến trích dẫn
- “Gemini is the first model capable of accepting 1 million tokens.” (bản dịch) «Gemini là đó đầu tiên model capable of accepting 1 million tokens.» Nhảy đến trích dẫn
Google DeepMind engineers — tokens và đó hardware ceiling (qua Google Blog)
- “measures how many tokens — the smallest building blocks, like part of a word, image or video — that the model can process at once.” (bản dịch) «measures cách nhiều tokens — đó smallest building chặn, như part of một word, image hoặc video — đó model có thể xử lý tại khi.» Nhảy đến trích dẫn
- “10 million tokens at once is already close to the thermal limit of our Tensor Processing Units.” (bản dịch) «10 million tokens tại khi là đã close để đó thermal limit of của chúng ta Tensor Processing Units.» — Nikolay Savinov, Research Scientist, Google DeepMind. Đọc bài đưa tin
Google — không ideal trang length, không requirement để fragment
- “There’s no ideal page length, and in the end, make pages for your audience, not just for generative AI search.” (bản dịch) «có không ideal trang length, và trong đó end, làm các trang cho của bạn audience, không chỉ cho generative AI tìm kiếm.» — Google Search Central, Optimizing của bạn website cho generative AI features. Đọc đó hướng dẫn
- “There’s no requirement to break your content into tiny pieces for AI to better understand it.” (bản dịch) «có không requirement để break nội dung của bạn vào tiny pieces cho AI để tốt hơn understand điều này.» — Google Search Central, giống nhau hướng dẫn. Đọc đó hướng dẫn
Microsoft — chunking để stay dưới token limits
- “Partitioning large documents into smaller chunks can help you stay under the maximum token input limits of chat completion and embedding models.” (bản dịch) «Partitioning lớn documents vào nhỏ hơn chunks có thể help bạn stay dưới đó maximum token input limits of chat completion và embedding models.» — Microsoft, Azure AI Tìm kiếm tài liệu. Đọc đó nguồn
Đó ngành vấn đề statement — llms.txt spec
- “Large language models increasingly rely on website information, but face a critical limitation: context windows are too small to handle most websites in their entirety.” (bản dịch) «Lớn language models increasingly rely on website information, nhưng face một cốt yếu limitation: context windows là cũng nhỏ để xử lý hầu hết websites trong của họ entirety.» — llmstxt.org spec. Nhảy đến trích dẫn
Tokens và context windows — bảng tra nhanh
Điều gì mỗi term là trong một line
| Term | Một-line definition |
|---|---|
| Token | Đó unit an LLM xử lý — một sub-word fragment, ~4 characters |
| Tokenization | Splitting text (và images/audio/video) vào tokens |
| Vocabulary | Đó đầy đủ set of tokens một model knows |
| Context window | Max tokens một model có thể hold tại khi (input + output) |
Token ↔ word rule of thumb (Google, Gemini)
| Đo lường | Approximate giá trị |
|---|---|
| 1 token | ~4 characters |
| 100 tokens | ~60–80 English words |
| 1 English word | ~1,3 tokens on average (một token là ~¾ of một word) |
(Caveat: averages chỉ — punctuation, code, và non-English text dùng hơn tokens theo word.)
Cách context windows grew
| Era | Approx. window |
|---|---|
| Sớm generative models | ~2 000 tokens |
| 2023-era models | 8K → 32K → 128K |
| Gemini (2024–26) | 1 000 000+ tokens |
| Experimental | ~10 000 000 (near đó TPU “thermal limit” (bản dịch) «thermal limit») |
Điều gì 1 million tokens có thể hold (Google các ví dụ)
- ~50 000 lines of code (80 chars/line)
- 5 năm of của bạn text messages
- 8 average-length English novels
- Transcripts of 200+ podcast episodes
Fast facts
- Context window = một shared budget cho hệ thống prompt + history + retrieved tài liệu + query + output.
- Nội dung bên ngoài đó window là invisible — đã biến mất, không dần forgotten.
- Bigger window ≠ tốt hơn các câu trả lời — “Lost in the Middle” (bản dịch) «Lost trong đó Middle»: bắt đầu/end beat đó middle.
- Trong tìm kiếm, retrieval chunks và selects trước đó model sees nội dung — một 1M window không có nghĩa là AI đọc của bạn toàn bộ site.
- Google: không ideal trang length, không requirement để fragment nội dung cho AI.
- Tokens là metered/billed (output thường costs hơn input) — đó economic reason các hệ thống retrieve selectively.
Đó mental models
1. Token = đó model unit of reading, không đó word. LLMs đọc trong sub-word fragments (~4 characters, ~¾ of một word). Khi bạn estimate cost hoặc budget, count tokens, không words — punctuation, code, và non-English text inflate đó ratio.
2. Đó context window là ngắn-term memory — một shared budget. Hệ thống prompt, history, retrieved documents, của bạn query, và đó model output all draw từ một pool. Bất cứ điều gì bên ngoài điều này là invisible, không fading. Nếu một model seems để “forget,” ask điều gì fell out of đó window.
3. Bigger window ≠ hơn understanding. Capacity và comprehension là khác nhau axes. “Lost in the Middle” (bản dịch) «Lost trong đó Middle» có nghĩa là đó bắt đầu và end of đó context nhận dùng tốt hơn đó middle — so lead với của bạn point, không chỉ thêm hơn.
4. Không trang fits — so retrieval selects đầu tiên. Đó toàn bộ reason AI tìm kiếm chunks và retrieves là đó một trang (let alone một site) không fit, hoặc không economical để feed, toàn bộ. Đó model không bao giờ sees trang của bạn; điều này sees đó passages một retrieval hệ thống chose. Optimize để là được chọn, không ingested toàn bộ.
5. Đó decision rule cho nội dung. không manually cut nội dung vào token-sized chặn (Google says bạn không cần để). Làm front-load các câu trả lời và ghi self-contained sections — vì cả hai retrieval và trong-context attention reward nội dung đó stands on của nó own và leads với của nó point.
Context-window mistakes
Treating words và tokens as interchangeable
Tokenization varies by model, language, punctuation, và code. Dùng đó tokenizer cho đó thực tế model khi một hard limit matters; một word-count ratio là chỉ planning hướng dẫn.
Filling đó entire advertised context window
Capacity không có nghĩa là mỗi token nhận equal attention hoặc đó output space là free. Reserve room cho instructions và đó câu trả lời, xóa duplicated material, và kiểm thử đó đầy đủ task thay vì celebrating một maximum input size.
Splitting nội dung tại một fixed character count
Blind cuts có thể tách biệt một heading, definition, bảng, hoặc qualifier từ đó passage điều này giải thích. Chunk on semantic boundaries và bảo toàn đủ local context cho mỗi unit để stand alone.
Đo lường trang sections trước chunking
Chạy này trong đó Chrome DevTools Console. Điều này các báo cáo characters và whitespace-split words cho mỗi headed section không có pretending những được tính equal model tokens:
const headings = [...document.querySelectorAll('main h2, main h3')];
console.table(headings.map((h, i) => { let text = ''; for (let n = h.nextElementSibling; n && !/^H[23]$/.test(n.tagName); n = n.nextElementSibling) text += ` ${n.innerText || ''}`; return { heading: h.textContent.trim(), characters: text.trim().length, words: text.trim() ? text.trim().split(/\s+/).length : 0, nextHeading: headings[i + 1]?.textContent.trim() || '' }; }));Dùng đó output để tìm oversized hoặc empty sections, thì chạy đó cuối text qua đó tokenizer cho đó model bạn sẽ thực ra dùng.
Tìm có khả năng semantic boundaries trong đơn giản text
Này regular expression matches Markdown second- và third-cấp độ headings và giữ đó heading text trong capture group 1:
^#{2,3}\s+(.+)$Dùng multiline chế độ. Headings là candidate boundaries, không phải là bảo đảm đó mỗi section là hoàn tất đủ để retrieve alone.
Các tài nguyên worth của bạn time
My related writing
- Điều gì We Thực ra Know Về Optimizing cho LLM Tìm kiếm — my Ahrefs piece on cách AI retrieval xử lý nội dung, including Dan Petrovic Chrome DocumentChunker research on cách các trang là chunked vào ~200-word passages trước họ là processed — đó token/context-window limit là chính xác vì sao đó chunking happens.
- Generative Engine Optimization — my rộng hơn take on optimizing cho một retrieval-và-generation tìm kiếm landscape shaped by những context limits.
My speaking
- GEO? AEO? LLMO? — my AI tìm kiếm webinar — nơi retrieval, chunking, và context limits fit vào AI tìm kiếm. (My standing disclaimer áp dụng: này là my understanding of những các hệ thống, không going để là 100% hoàn tất hoặc chính xác.)
Chính thức
- Understand và count tokens (Gemini API) và Dài context (Gemini API) — Google own definitions và đó million-token các ví dụ.
- Optimizing của bạn website cho generative AI features (Google Search Central) — “no ideal page length,” (bản dịch) «không ideal trang length,» “no requirement to break your content into tiny pieces.” (bản dịch) «không requirement để break nội dung của bạn vào tiny pieces.»
Từ khoảng đó ngành
- Điều gì là một dài context window? Google DeepMind engineers giải thích — Nikolay Savinov on đó TPU thermal ceiling, plus Machel Reid và Denis Teplyashin on điều gì dài context unlocks.
- Chunk lớn documents cho vector tìm kiếm (Microsoft Azure AI Tìm kiếm) — chunking để stay dưới maximum token input limits.
- Đó llms.txt spec — đó “context windows are too small to handle most websites in their entirety” (bản dịch) «context windows là cũng nhỏ để xử lý hầu hết websites trong của họ entirety» vấn đề statement, và đó format được xây dựng khoảng điều này.
- SEO Chunk Optimization là Overrated (Despina Gavoyannis, Ahrefs) — “You can’t control how Google, ChatGPT, or Perplexity chunk your content. Their pipelines change based on cost, model, and context.” (bản dịch) «Bạn không thể control cách Google, ChatGPT, hoặc Perplexity chunk nội dung của bạn. Của họ pipelines thay đổi dựa trên cost, model, và context.»
- Chunked, Retrieved, Synthesized — Không Được xếp hạng (Duane Forrester) — “If traditional SEO optimized for clicks, GenAI systems optimize for chunks… Structure still wins.” (bản dịch) «Nếu truyền thống SEO optimized cho clicks, GenAI các hệ thống optimize cho chunks… Structure vẫn wins.»
- Điều gì là một context window? (IBM) — một neutral, thorough reference explainer.
Số liệu worth citing
- ~4 characters theo token; 100 tokens ≈ 60–80 English words — Google own rule of thumb cho Gemini models, đó hầu hết citable token-để-word conversion. Nguồn
- 1 000 000 tokens — Gemini là đó đầu tiên model capable of accepting 1 million tokens trong một single window, lên từ ~2K trong sớm generative models. Nguồn
- ~10 000 000 tokens = near đó TPU thermal limit — Google DeepMind Nikolay Savinov on đó hiện tại physical ceiling cho context-window size. Nguồn
- 1 million tokens ≈ 8 novels / 200+ podcast transcripts / 50 000 lines of code — Google concrete sizing cho điều gì một million-token window có thể hold. Nguồn
- Token-based billing (output thường costs hơn input) — theo Google, đó cost of một Gemini API call là determined trong part by input và output token được tính, đó economic reason AI các hệ thống retrieve selectively thay vì feed toàn bộ các trang. Nguồn
Tự kiểm tra: Tokens và Context Windows
Five nhanh các câu hỏi on cách LLMs đọc text và cách nhiều they có thể hold tại khi. Pick an câu trả lời cho mỗi, thì kiểm tra.
Nhật ký thay đổi
Đã cập nhật 8 thg 8, 2026.
Tóm tắt biên tập và chi tiết thay đổi đã ghi nhận.Chi tiết thay đổi
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
Không thể so sánh đầy đủ — không có bản lưu trước đó cho lần sửa đổi này.
Đã cập nhật 19 thg 7, 2026.
Tóm tắt biên tập và chi tiết thay đổi đã ghi nhận.Chi tiết thay đổi
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
Không thể so sánh đầy đủ — không có bản lưu trước đó cho lần sửa đổi này.