Vibe Xếp hạng
Điều gì 'vibe xếp hạng' thực sự có nghĩa là — LLM-based reranking, đó hai-stage retrieval pipeline, điều gì Google và Bing thực ra làm, và cách optimize cho điều này.
Ngôn ngữ
"Vibe ranking" _(bản dịch)_ «Vibe xếp hạng» là practitioner shorthand — không an chính thức Google hoặc Bing term — cho LLM-based reranking: dùng một lớn language model để reorder retrieved tìm kiếm candidates by holistic, semantic judgment thay vì từ khóa overlap hoặc thô vector similarity. Đó real mechanism behind điều này là đó hai-stage retrieval pipeline (retrieve nhiều candidates cheaply, thì rerank một nhỏ set với an expensive model) và three reranking paradigms (pointwise, pairwise, listwise). Google DOJ-disclosed RankEmbed và pairwise patent, plus Bing Web IQ và Google được ghi lại Passage xếp hạng hệ thống, xác nhận LLMs và passage-cấp độ scoring là hiện tại deep trong xếp hạng — though neither vendor tiết lộ đó scoring function. Đó SEO takeaway: clear, có thẩm quyền, well-structured, factually grounded prose là defensible practice dưới an LLM reranker, không một guaranteed win.
TL;DR — “Vibe ranking” (bản dịch) «Vibe xếp hạng» là một nickname cho một real điều: AI các công cụ tìm kiếm dùng một lớn language model để đọc một set of kết quả tìm kiếm và reorder them by cách good they thực ra là — clear, expert, on-topic — thay vì chỉ counting từ khóa. đây là không an chính thức Google hoặc Bing term. Đó name borrows từ “vibe coding.” (bản dịch) «vibe coding.» Đó practical lesson: ghi clear, trustworthy passages đó câu trả lời đó câu hỏi trực tiếp.
Điều gì “vibe ranking” (bản dịch) «vibe xếp hạng» có nghĩa là
“Vibe ranking” (bản dịch) «Vibe xếp hạng» là an informal label ở đây, không một standardized research hoặc vendor term. Evidence for this claim Vibe ranking is used in this article as informal editorial shorthand, not as a standardized research or vendor term. Scope: Terminology boundary for this article; the cited research supports retrieve-then-rerank architectures, not the phrase vibe ranking. Confidence: medium · Verified: Reimers and Gurevych: Sentence-BERT Đó underlying retrieve-thì-rerank pattern là established: một fast retriever produces candidates và một hơn expensive relevance model có thể reorder them. Evidence for this claim Sentence-BERT research describes efficient bi-encoder retrieval and cross-encoder-style pair scoring, supporting a retrieve-then-rerank pattern. Scope: The paper's evaluated models and datasets; production ranking stacks may use different candidates, models, and signals. Confidence: high · Verified: Reimers and Gurevych: Sentence-BERT
Đầu tiên, đó honest part: “vibe ranking” (bản dịch) «vibe xếp hạng» không phải một term Google hoặc Bing dùng. Bạn sẽ không tìm điều này trong của họ tài liệu. đây là shorthand đó SEO mọi người đã bắt đầu dùng — borrowed từ “vibe coding,” (bản dịch) «vibe coding,» đó phrase Andrej Karpathy coined trong sớm 2025 cho loosely steering an AI by feel thay vì writing mỗi line yourself.
Đó “real name” cho đó điều này mô tả là LLM reranking. Ở đây đó ý tưởng.
Khi bạn ask an AI tìm kiếm tool một câu hỏi, điều này không chỉ grab đó single best trang. Điều này hoạt động trong hai steps:
- Retrieve một pile of candidates. MỘT fast, cheap hệ thống pulls lại maybe hundreds of có thể passages — bất cứ điều gì đó looks khoảng relevant.
- Rerank đó best một vài. MỘT chậm hơn, smarter model thì đọc một nhỏ set of những candidates và decides mà ones thực ra câu trả lời đó câu hỏi best.
Đó second step là nơi đó “vibe” xuất hiện trong. Thay vì chỉ counting cách nhiều times của bạn từ khóa xuất hiện, một lớn language model đọc đó passage đó way một human reviewer sẽ — và judges điều như: Là này clear? Làm điều này thực ra câu trả lời đó câu hỏi? Làm đó nguồn seem credible? Là điều này hoàn tất?
Vì sao này matters cho nội dung của bạn
Old-school SEO advice đã là đầy đủ of tricks cho matching từ khóa. LLM reranking cares một lot ít hơn về đó và một lot hơn về liệu của bạn writing là genuinely good.
Vài điều đó help:
- Câu trả lời đó câu hỏi trực tiếp. không bury đó câu trả lời dưới throat-clearing.
- Là clear và well-organized. Ngắn paragraphs, logical order, đơn giản language.
- Là trustworthy. Nhận facts right, và làm them checkable.
- Là hoàn tất. Cover đó toàn bộ câu hỏi, không chỉ một slice of điều này.
có một hơn shift worth knowing: AI tìm kiếm increasingly scores passages, không toàn bộ các trang. So một mạnh paragraph on một topic có thể nhận pulled vào an câu trả lời ngay cả khi đó rest of trang của bạn là về điều gì đó khác — và một yếu paragraph có thể lose, ngay cả on an nếu không great trang.
Muốn đó real mechanics — đó hai-stage pipeline, đó three ways LLMs rerank, và điều gì Google và Bing có thực ra confirmed? Chuyển để đó Advanced tab.
TL;DR — “Vibe ranking” (bản dịch) «Vibe xếp hạng» là informal practitioner shorthand (by analogy với “vibe coding”) cho LLM-based reranking — không an chính thức Google hoặc Bing term. Đó real mechanism là đó hai-stage retrieval pipeline: cheap đầu tiên-stage retrieval (BM25 / bi-encoders) cho recall, thì an expensive second-stage reranker cho precision. LLM rerankers come trong three flavors — pointwise, pairwise, listwise — với pairwise và listwise outperforming pointwise trong đó research. Google DOJ-disclosed RankEmbed (LLM-trained) và một pairwise patent, plus Bing Web IQ passage-cấp độ evidence objects, xác nhận LLMs là deep trong xếp hạng — though neither xác nhận một discrete “LLM reranker” (bản dịch) «LLM reranker» stage by đó name. Đó SEO consequence là đó passage-cấp độ shift: của bạn paragraph competes so với các đối thủ’ paragraphs on holistic quality.
MỘT note on đó term trước we bắt đầu
Này bài viết dùng “vibe ranking” (bản dịch) «vibe xếp hạng» as editorial shorthand cho learned reranking dựa trên rộng relevance các tín hiệu; readers không nên treat điều này as một named tìm kiếm-hệ thống tiêu chuẩn. Evidence for this claim Vibe ranking is used in this article as informal editorial shorthand, not as a standardized research or vendor term. Scope: Terminology boundary for this article; the cited research supports retrieve-then-rerank architectures, not the phrase vibe ranking. Confidence: medium · Verified: Reimers and Gurevych: Sentence-BERT Research on sentence embeddings và cross-encoders hỗ trợ đó rộng hơn hai-stage retrieval và reranking architecture. Evidence for this claim Sentence-BERT research describes efficient bi-encoder retrieval and cross-encoder-style pair scoring, supporting a retrieve-then-rerank pattern. Scope: The paper's evaluated models and datasets; production ranking stacks may use different candidates, models, and signals. Confidence: high · Verified: Reimers and Gurevych: Sentence-BERT
Let me là upfront, vì I’d rather coin một hữu ích frame honestly hơn pretend đây là chính thức: “vibe ranking” (bản dịch) «vibe xếp hạng» không phải một term Google, Bing, hoặc bất kỳ công cụ tìm kiếm dùng. Điều này không xuất hiện trong của họ tài liệu, patents, hoặc engineer statements. đây là emerging practitioner shorthand, gần như certainly borrowed by analogy từ vibe coding (coined by Andrej Karpathy trong February 2025, popular đủ để become Collins Dictionary Word of đó Năm cho 2025).
I’m dùng điều này anyway vì đây là một good metaphor cho một real, well-được ghi lại mechanism — đó formal name cho mà là LLM-based reranking (cũng: neural reranking, passage reranking, pairwise/listwise xếp hạng). Đó “vibe” captures đó key khác biệt: an LLM đọc một passage holistically, như một human judge, rather hơn computing một similarity score từ surface features. Chỉ không quote điều này lại để me as điều gì đó Google đã nói. Điều này không.
Đó hai-stage retrieval pipeline
Này là đó part đó là đã đúng trong thời gian dài, well trước LLMs. Retrieval chạy trong hai stages vì bạn không thể afford để làm đó expensive điều để mọi thứ:
- Stage 1 — retrieval (recall). MỘT fast, cheap phương thức — từ khóa matching với BM25, hoặc một bi-encoder đó turns các truy vấn và documents vào vectors — pulls lại một lớn candidate set (top 100–1000) trong milliseconds. Cheap, nhưng imprecise.
- Stage 2 — reranking (precision). MỘT chậm hơn, hơn chính xác model re-scores đó nhỏ candidate set để tìm đó genuinely best handful. Này được dùng để là một cross-encoder; hiện tại điều này có thể là an LLM.
Vì sao bother với hai stages? Vì đó precise phương thức không quy mô. As Pinecone diễn đạt điều này, đang chạy BERT over 40M records on một GPU sẽ take “more than 50 hours,” (bản dịch) «hơn 50 hours,» versus dưới 100ms với vector tìm kiếm alone. Bạn retrieve broadly với đó cheap tool, thì spend đó expensive compute chỉ on đó survivors. Và cho AI các câu trả lời cụ thể, điều này matters twice over: “LLM recall degrades as we put more tokens in the context window.” (bản dịch) «LLM recall degrades as we put hơn tokens trong đó context window.» Feeding đó generator đó best 5 passages beats dumping 200 mediocre ones vào đó prompt.
Bi-encoder so với cross-encoder so với LLM, since này là đó toàn bộ spectrum:
- Bi-encoders embed đó query và đó document riêng, so document vectors có thể là precomputed. Fast tại query time, nhưng they lose đó query↔document interaction.
- Cross-encoders chạy đó query và document together qua một transformer. Far hơn chính xác, far chậm hơn — bạn không thể precompute.
- LLM rerankers là đó cross-encoder ý tưởng taken để đó extreme: đầy đủ generative understanding, reasoning về relevance, quality, và tính đầy đủ — tại đó highest cost và latency of đó three.
Đó three reranking paradigms
Này là đó cốt lõi of “how LLMs actually rerank,” (bản dịch) «cách LLMs thực ra rerank,» và đây là nơi đó academic literature là solid. Có three ways để ask an LLM để xếp hạng:
- Pointwise — score mỗi passage independently (“how relevant is this passage, 0–1?” (bản dịch) «cách relevant là này passage, 0–1?»). Đơn giản, parallelizable, nhưng LLMs là bad tại producing calibrated absolute scores, và đây là đó hầu hết expensive theo unit of quality.
- Pairwise — cho thấy đó model hai passages và ask mà là tốt hơn cho đó query. Chạy twice với đó order đã đổi để cancel position bias. Này plays để một real LLM strength: as đó Pairwise Xếp hạng Prompting (PRP) authors put điều này, LLMs có “a sense of pairwise relative comparisons, which is much simpler than requiring calibrated pointwise relevance estimation.” (bản dịch) «một hợp lý of pairwise relative comparisons, mà là nhiều simpler hơn requiring calibrated pointwise relevance estimation.»
- Listwise — hand đó model đó toàn bộ set và ask điều này để output một được xếp hạng permutation tại khi. Này là đó hầu hết “vibe-như” approach: một holistic judgment over đó set. RankGPT làm này với một sliding window over đó candidates.
Đó research consensus là đó pointwise là đó yếu option và pairwise/listwise win. ZeroEntropy benchmarking là blunt: “Pointwise LLM reranking is almost never worth it — 10x the cost and lower accuracy than specialized rerankers.” (bản dịch) «Pointwise LLM reranking là gần như không bao giờ worth điều này — 10x đó cost và thấp hơn độ chính xác hơn specialized rerankers.» Listwise LLM reranking có thể beat specialized cross-encoders, nhưng tại một steep latency/cost premium (của họ numbers: một listwise LLM tại NDCG@10 0,78 / 420ms / ~9x cost so với một specialized reranker tại 0,74 / 12ms).
Hai landmark kết quả worth knowing:
- RankGPT (Sun et al., EMNLP 2023, “Is ChatGPT Good at Search?” (bản dịch) «Là ChatGPT Good tại Tìm kiếm?») showed đó “properly instructed LLMs can deliver competitive, even superior results to state-of-the-art supervised methods” (bản dịch) «properly instructed LLMs có thể deliver competitive, ngay cả superior kết quả để state-of-đó-art supervised các phương thức» — zero-shot. GPT-4 produced “remarkable results” (bản dịch) «remarkable kết quả» on TREC benchmarks, và đó authors distilled đó xếp hạng ability vào một nhiều nhỏ hơn model đó beat một 3B supervised baseline.
- PRP (Qin et al.) showed FLAN-UL2 (20B) với pairwise prompting beating InstructGPT (175B) by >10% on TREC-DL2019 — và đang far hơn robust để input order hơn listwise RankGPT, mà collapsed từ NDCG@10 65,80 để 32,77 khi đó input order đã là reversed.
Điều gì Google và Bing là thực ra đang làm
Ở đây nơi I tách biệt confirmed từ speculation, vì đó khoảng trống matters.
Google — confirmed LLMs là trong xếp hạng; “reranker stage” (bản dịch) «reranker stage» unconfirmed.
- RankEmbed là, theo DOJ antitrust testimony từ Pandu Nayak (disclosed sớm 2025, reported qua Search Engine Land), một “primary Google signal, trained with Large Language Models.” (bản dịch) «chính Google tín hiệu, trained với Lớn Language Models.» đây là một dual-encoder đó maps các truy vấn và các trang vào an embedding space và ranks by distance. Note: này là an LLM-trained retrieval/tín hiệu xếp hạng, không một discrete hai-stage “reranker” by đó academic definition. (DOJ trial testimony — court-disclosed, không một Google publication.)
- Đó pairwise patent (US20250124067A1, “Method for Text Ranking with Pairwise Ranking Prompting,” (bản dịch) «Phương thức cho Text Xếp hạng với Pairwise Xếp hạng Prompting,» Google LLC, filed Oct 2024, published Apr 2025) mô tả chính xác đó pairwise approach trên: an LLM compares passage pairs, chạy twice cho position bias, được tổng hợp qua all-pairs, sliding window, hoặc sorting. MỘT patent là không một deployed feature — treat này as architecture Google có worked on, không confirmed production behavior.
- MỘT passage-cấp độ xếp hạng hệ thống là officially được ghi lại — chỉ không đó quote này bài viết được dùng để cite. Google own hướng dẫn để Tìm kiếm xếp hạng các hệ thống lists một “Passage ranking system,” (bản dịch) «Passage xếp hạng hệ thống,» described ở đó as an AI hệ thống “we use to identify individual sections or ‘passages’ of a web page to better understand how relevant a page is to a search.” (bản dịch) «we dùng để identify riêng lẻ sections hoặc ‘passages’ of một web trang để tốt hơn understand cách relevant một trang là để một tìm kiếm.» đó là real, hiện tại sub-document granularity — nhưng Google text ties điều này để chung Tìm kiếm relevance, không để Gemini hoặc AI Overviews by name. (I checked đó AI Optimization Hướng dẫn trực tiếp by thô fetch, including một Wayback snapshot từ đó day này brief đã là researched — điều này contains không mention of “Gemini” hoặc “passage” tại all. An trước đó draft of này bài viết quoted một “Gemini… passage indexing” (bản dịch) «Gemini… passage lập chỉ mục» line as nếu điều này đã là on đó trang; điều này không và không bao giờ đã là. Đó đã là my lỗi, và I’ve corrected điều này ở đây.)
Điều gì Google có không published: bất kỳ mô tả of một distinct LLM reranking stage bên trong AI Overviews — không cross-encoder, không pairwise step, không listwise truyền confirmed trong đó trực tiếp pipeline, và không có gì đó names passage xếp hạng as part of đó pipeline cụ thể.
Bing — đó clearest chính thức xác nhận of passage-cấp độ, LLM-aware scoring.
Microsoft Web IQ (announced June 2026, Knut Risvik) là đó strongest on-record tín hiệu từ bất kỳ major engine. Của nó model layer bao gồm “our best-in-class embedding model, which defines how information is projected into a space where semantic similarity becomes computationally tractable,” (bản dịch) «của chúng ta best-trong-class embedding model, mà defines cách information là projected vào một space nơi semantic similarity becomes computationally tractable,» alongside tách biệt “models that are optimized for content understanding and ranking, trained not for isolated metrics but for how their outputs are used inside LLM-driven reasoning.” (bản dịch) «models đó là optimized cho nội dung understanding và xếp hạng, trained không cho isolated các chỉ số nhưng cho cách của họ outputs là dùng bên trong LLM-driven reasoning.» Đó cuối cùng phrase là về as close as anyone có come để officially confirming an LLM-aware reranking layer. Web IQ trả về passages và structured evidence objects, không đầy đủ documents, on đó principle of “fewer tokens in, better answers out, lower cost per call.” (bản dịch) «ít hơn tokens trong, tốt hơn các câu trả lời out, thấp hơn cost theo call.»
Bing trước đó “Evolving role of the index” (bản dịch) «Evolving role of đó chỉ mục» post (Có thể 2026) frames đó shift well: “Search indexing was built to help humans decide what to read. Grounding indexing is being built to help AI systems decide what to say.” (bản dịch) «Tìm kiếm lập chỉ mục đã là được xây dựng để help humans decide điều cần đọc. Grounding lập chỉ mục là đang được xây dựng để help AI các hệ thống decide điều cần chẳng hạn.»
Đó passage-cấp độ shift — đó part đó thực ra thay đổi của bạn SEO
Nếu bạn take một operational điều từ này bài viết, take này: đó unit of competition là moving từ đó document để đó passage. Google được ghi lại passage xếp hạng hệ thống và Bing evidence objects cả hai score discrete chunks. Combine đó với pairwise reranking, và của bạn riêng lẻ paragraph là đang compared head-để-head so với một đối thủ paragraph on đó giống nhau sub-topic.
Đó Search Engine Land cách diễn đạt of đó Google patent captures đó consequence: nội dung “doesn’t compete in isolation but undergoes relative evaluation against all surviving candidates.” (bản dịch) «không compete trong isolation nhưng undergoes relative evaluation so với all surviving candidates.» MỘT great trang với một yếu paragraph on một sub-topic có thể lose đó sub-topic để một weaker trang với một mạnh paragraph. Này là cũng vì sao đó modern agentic loop matters hơn bất kỳ single xếp hạng moment: AI tìm kiếm fans một câu hỏi out vào nhiều sub-các truy vấn, retrieves và reranks cho mỗi, và chạy một critic truyền. Nội dung của bạn có để survive repeatedly, không xếp hạng #1 khi.
Worth đang precise về: clear, self-contained passages là một defensible usability và retrieval practice regardless — họ là easier cho bất kỳ retriever để match và bất kỳ reader để dùng. Nhưng không nguồn ở đây proves một well-được viết passage universally wins bên trong an undisclosed reranker. Google và Bing xác nhận đó architecture operates tại passage granularity; neither publishes đó scoring function. Treat “write good passages” (bản dịch) «ghi good passages» as sound practice, không một guaranteed xếp hạng lever.
Này connects trực tiếp để đó rest of cách AI các câu trả lời nhận được xây dựng — passage xếp hạng, chunking, embeddings, vector tìm kiếm, và RAG/grounding là all upstream và downstream of đó reranking step.
Điều gì LLM rerankers favor (và đó honest caveat)
Synthesizing đó PRP research và practitioner analysis of đó pairwise patent, đó nội dung đó wins relative comparisons tends để share những traits:
- Trực tiếp intent match — các câu trả lời đó query không có tangential padding.
- Semantic tính đầy đủ — addresses all components of đó câu hỏi.
- Factual density với clear provenance — checkable claims, có thể gán sources.
- Clear, logically organized writing — structure đó model có thể follow.
- Có thẩm quyền, trustworthy tone — và genuine subject expertise behind điều này.
Luca Tagliaferro đọc on đó patent sums lên đó mental shift: xếp hạng moves từ “absolute, deterministic relevance to relative, model-mediated probabilistic relevance.” (bản dịch) «absolute, deterministic relevance để relative, model-mediated probabilistic relevance.»
Đó caveat I sẽ không skip: có không single universal “vibe.” Research diagnosing LLM rerankers dưới fixed evidence pools được tìm thấy they exhibit “model-specific, non-uniform behavior that cannot be reduced to a single recognizable optimization strategy.” (bản dịch) «model-cụ thể, non-uniform behavior đó không thể là reduced để một single recognizable optimization strategy.» Llama implicitly diversifies; GPT increases redundancy; Qwen sits trong giữa. họ là cũng không lexical matchers — BM25 agreement đã là yếu (τ từ ~0,19 để ~0,41). So “optimize for the AI’s vibe” (bản dịch) «optimize cho đó AI vibe» là an oversimplification: bạn optimize cho genuinely clear, hoàn tất, well-sourced passages, vì đó là điều gì holds lên trên models — không cho một model quirks.
AI summary
MỘT condensed take on đó Advanced version:
- “Vibe ranking” (bản dịch) «Vibe xếp hạng» là không chính thức. đây là practitioner shorthand (borrowed từ “vibe coding,” (bản dịch) «vibe coding,» Karpathy, Feb 2025) cho LLM-based reranking. Không công cụ tìm kiếm dùng đó term.
- Đó real mechanism là hai-stage retrieval: cheap đầu tiên-stage retrieval (BM25 / bi-encoders) cho recall, thì an expensive second-stage reranker cho precision. LLM rerankers là đó newest second stage, sau cross-encoders.
- Three paradigms: pointwise (score mỗi — weakest, priciest), pairwise (so sánh hai, đổi để cancel position bias — an LLM strength), listwise (xếp hạng đó toàn bộ set tại khi — đó “vibe” approach). Pairwise/listwise beat pointwise.
- Landmark kết quả: RankGPT (GPT-4 competitive/superior zero-shot on TREC); PRP (20B model beating một 175B một by >10%, và far hơn order-robust hơn listwise).
- Google: RankEmbed là LLM-trained (DOJ testimony) nhưng không một discrete reranker; một pairwise patent tồn tại (filed Oct 2024) nhưng một patent ≠ production; một được ghi lại “Passage ranking system” (bản dịch) «Passage xếp hạng hệ thống» tồn tại (Google xếp hạng-các hệ thống hướng dẫn), nhưng đây là không named as part of AI Overviews hoặc tied để Gemini trong Google own text.
- Bing: Web IQ dùng models “trained not for isolated metrics but for how their outputs are used inside LLM-driven reasoning” (bản dịch) «trained không cho isolated các chỉ số nhưng cho cách của họ outputs là dùng bên trong LLM-driven reasoning» và trả về passage-cấp độ evidence objects — đó clearest chính thức tín hiệu of LLM-aware scoring.
- SEO consequence — đó passage-cấp độ shift: của bạn paragraph competes head-để-head so với các đối thủ’ paragraphs. Win với trực tiếp intent match, tính đầy đủ, factual provenance, clear structure, và authority.
- Caveat: không universal “vibe” — rerankers (Llama/GPT/Qwen) behave differently; họ là không lexical matchers. Optimize cho genuinely good passages, không một model.
Reranking paradigms — bảng tra nhanh
Đó hai-stage pipeline tại một glance
| Stage | Phương thức | Job | Speed | Cost |
|---|---|---|---|---|
| 1 — Retrieval | BM25, bi-encoder | Recall: pull top 100–1000 candidates | Milliseconds | Cheap |
| 2 — Reranking | Cross-encoder, LLM reranker | Precision: re-score đó survivors | Chậm | Expensive |
Đó three LLM reranking paradigms
| Paradigm | Cách nó hoạt động | Strength | Weakness |
|---|---|---|---|
| Pointwise | Score mỗi passage independently | Đơn giản, parallel | Poor calibration; ~10x cost, thường thấp hơn độ chính xác hơn specialized rerankers |
| Pairwise | So sánh hai passages; đổi order để cancel position bias | Plays để LLMs’ relative-so sánh strength; order-robust | O(N²) naively (mitigated by sliding window / sorting) |
| Listwise | Xếp hạng đó toàn bộ set tại khi (sliding window) | Hầu hết holistic — đó “vibe” approach; có thể beat specialized rerankers | Fragile để input order; cao latency/cost |
Encoder spectrum
| Loại | Query + doc processed | Precompute? | Speed | Độ chính xác |
|---|---|---|---|---|
| Bi-encoder | Riêng | Có | Fast | Thấp hơn (không interaction) |
| Cross-encoder | Together | Không | Chậm | Cao |
| LLM reranker | Together, generatively | Không | Slowest | Highest (reasons về quality) |
Fast facts
- “Vibe ranking” (bản dịch) «Vibe xếp hạng» = informal term; formal name là LLM-based reranking.
- Pointwise LLM reranking: “almost never worth it” (bản dịch) «gần như không bao giờ worth điều này» (ZeroEntropy).
- PRP (20B) beat InstructGPT (175B) by >10% on TREC-DL2019.
- Google RankEmbed = LLM-trained dual encoder (DOJ testimony), không một discrete reranker.
- Bing Web IQ trả về passages/evidence objects, không đầy đủ documents.
- LLM rerankers là không lexical matchers (BM25 agreement τ ≈ 0,19–0,41).
Đó mental models
1. Retrieve rộng, rerank hẹp. Đầu tiên stage maximizes recall cheaply; second stage maximizes precision expensively. Bạn không thể chạy đó expensive model on mọi thứ — đó là đó entire reason hai stages exist. Khi an AI câu trả lời misses nội dung của bạn, ask mà stage dropped điều này: không bao giờ retrieved, hoặc retrieved nhưng reranked away?
2. Đó three paradigms — pointwise / pairwise / listwise.
- Pointwise: absolute score theo passage (weakest cho LLMs — poor calibration).
- Pairwise: head-để-head, đổi order để cancel position bias (an LLM strength).
- Listwise: xếp hạng đó toàn bộ set tại khi (đó holistic “vibe” — nhưng order-fragile). Default assumption cho production-grade các hệ thống: pairwise hoặc listwise, không pointwise.
3. Đó unit of competition là đó passage, không đó trang. Google được ghi lại passage xếp hạng hệ thống và Bing evidence objects có nghĩa là của bạn paragraph competes so với một đối thủ paragraph. MỘT mạnh trang với một yếu sub-topic paragraph loses đó sub-topic. Optimize passages, không chỉ các trang — nhưng note neither vendor tiết lộ đó thực tế scoring function, so treat này as sound practice, không một guaranteed win.
4. Relative, không absolute. Pairwise reranking làm relevance comparative: “of these two, which is better for this query?” (bản dịch) «of những hai, mà là tốt hơn cho này query?» bạn là không hitting một fixed bar — bạn là beating đó other survivors. Đó shift là từ “absolute, deterministic relevance to relative, model-mediated probabilistic relevance.” (bản dịch) «absolute, deterministic relevance để relative, model-mediated probabilistic relevance.»
5. Survive đó loop, không một single xếp hạng. Agentic AI tìm kiếm fans một câu hỏi vào nhiều sub-các truy vấn, reranks cho mỗi, và chạy một critic/reflection truyền (sufficiency, contradiction, freshness, nguồn diversity). Nội dung phải survive repeatedly trên retrieval và reflection — citation tracking chỉ sees cuối-stage survivors và có thể miss hầu hết of đó pipeline.
6. có không single “vibe.” Rerankers behave differently by model (Llama diversifies; GPT adds redundancy). không chase một model quirks. Optimize cho điều gì holds trên all of them: clear, hoàn tất, well-sourced, trực tiếp-on-intent passages.
Tài liệu chính thức
Chính-nguồn material on cách AI features retrieve và xếp hạng. Note lên front: “vibe ranking” (bản dịch) «vibe xếp hạng» xuất hiện trong none of những — they mô tả LLM-trained xếp hạng, passage-cấp độ scoring, và grounding, mà là điều gì đó term informally points tại.
- AI features và của bạn website (AI Optimization Hướng dẫn) — cách AI Overviews/AI Chế độ rely on cốt lõi Tìm kiếm xếp hạng, đó grounding/RAG definition, query fan-out, và đó passage-lập chỉ mục mention.
- Cách Google Search Hoạt động — đó crawl → chỉ mục → serve pipeline reranking sits bên trong.
- Patent US20250124067A1 — Phương thức cho Text Xếp hạng với Pairwise Xếp hạng Prompting — Google LLC, filed Oct 2024, published Apr 2025. Pairwise LLM passage so sánh. (MỘT patent, không một confirmed trực tiếp feature.)
- Re-xếp hạng (Google ML / Khuyến nghị Các hệ thống) — Google chung cách diễn đạt of một re-xếp hạng stage.
- MỘT hướng dẫn để Google Search xếp hạng các hệ thống — lists đó Passage xếp hạng hệ thống: “an AI system we use to identify individual sections or ‘passages’ of a web page to better understand how relevant a page is to a search.” (bản dịch) «an AI hệ thống we dùng để identify riêng lẻ sections hoặc ‘passages’ of một web trang để tốt hơn understand cách relevant một trang là để một tìm kiếm.» Này là đó real, hiện tại, chính thức passage-cấp độ hệ thống — đây là không tied để Gemini hoặc AI Overviews by name trong Google own text.
Bing / Microsoft
- Announcing Microsoft Web IQ (June 2026) — passage-cấp độ evidence objects; nội dung-understanding và xếp hạng models “trained not for isolated metrics but for how their outputs are used inside LLM-driven reasoning.” (bản dịch) «trained không cho isolated các chỉ số nhưng cho cách của họ outputs là dùng bên trong LLM-driven reasoning.»
- Evolving role of đó chỉ mục (Có thể 2026) — tìm kiếm lập chỉ mục so với grounding lập chỉ mục; factual fidelity, attribution, freshness, conflict detection.
- Introducing Deep Tìm kiếm (Dec 2023) — GPT-4 powered query expansion và relevance scoring over ~10x đó thông thường trang volume.
Quotes từ đó nguồn
On-đó-record statements. Nơi một deep link là khả dụng điều này jumps để đó quoted passage. Court-disclosed material là flagged riêng — điều này xuất hiện từ DOJ antitrust testimony, không từ một Google publication.
Google — grounding & AI features (tài liệu chính thức)
- “A technique (also known as grounding) used to improve the quality, accuracy, and freshness of AI responses by relying on our core Search ranking systems to retrieve relevant, up-to-date web pages from our Search index.” (bản dịch) «MỘT technique (cũng known as grounding) được dùng để improve đó quality, độ chính xác, và freshness of AI các phản hồi by relying on của chúng ta cốt lõi Tìm kiếm xếp hạng các hệ thống để retrieve relevant, lên-để-date web các trang từ của chúng ta Tìm kiếm chỉ mục.» — Google Search Central, AI Optimization Hướng dẫn.
- “Passage ranking is an AI system we use to identify individual sections or ‘passages’ of a web page to better understand how relevant a page is to a search.” (bản dịch) «Passage xếp hạng là an AI hệ thống we dùng để identify riêng lẻ sections hoặc ‘passages’ of một web trang để tốt hơn understand cách relevant một trang là để một tìm kiếm.» — Google Search Central, “A guide to Google Search ranking systems.” (bản dịch) «MỘT hướng dẫn để Google Search xếp hạng các hệ thống.» (Real, hiện tại, sub-document granularity — nhưng Google own text không tie này để Gemini hoặc AI Overviews by name. An trước đó draft of này bài viết attributed một “Gemini… passage indexing” (bản dịch) «Gemini… passage lập chỉ mục» quote để đó AI Optimization Hướng dẫn; đó trang contains không mention of “Gemini” hoặc “passage” — confirmed by trực tiếp fetch và một Wayback snapshot từ đó brief own research date. Đó quote đã là fabricated và đã được đã xóa.) Nguồn
Google — RankEmbed & pairwise xếp hạng (court-disclosed / patent)
- RankEmbed là described as một “primary Google signal, trained with Large Language Models” (bản dịch) «chính Google tín hiệu, trained với Lớn Language Models» — một dual encoder đó ranks qua distance trong an embedding space. Nguồn: DOJ antitrust trial testimony từ Google Pandu Nayak (disclosed sớm 2025), reported by Search Engine Land — court-disclosed document, không một Google publication. Đọc bài đưa tin
- Google patent mô tả các hệ thống đó “rank passages by having an LLM perform pairwise comparisons — of these two passages, which is better for this query?” (bản dịch) «xếp hạng passages by có an LLM perform pairwise comparisons — of những hai passages, mà là tốt hơn cho này query?» — meaning nội dung “doesn’t compete in isolation but undergoes relative evaluation against all surviving candidates.” (bản dịch) «không compete trong isolation nhưng undergoes relative evaluation so với all surviving candidates.» Này là một filed patent (US20250124067A1), không xác nhận of một deployed feature. Patent · SEL analysis
Bing / Microsoft — Web IQ & đó chỉ mục (chính thức blog)
- “Models that are optimized for content understanding and ranking, trained not for isolated metrics but for how their outputs are used inside LLM-driven reasoning.” (bản dịch) «Models đó là optimized cho nội dung understanding và xếp hạng, trained không cho isolated các chỉ số nhưng cho cách của họ outputs là dùng bên trong LLM-driven reasoning.» — Knut Risvik, Microsoft, Web IQ announcement (June 2026). Đó clearest chính thức tín hiệu of LLM-aware passage scoring. Nguồn
- “Fewer tokens in, better answers out, lower cost per call.” (bản dịch) «Ít hơn tokens trong, tốt hơn các câu trả lời out, thấp hơn cost theo call.» — Microsoft Web IQ design principle (passage-cấp độ evidence over đầy đủ documents). Nguồn
- “Search indexing was built to help humans decide what to read. Grounding indexing is being built to help AI systems decide what to say.” (bản dịch) «Tìm kiếm lập chỉ mục đã là được xây dựng để help humans decide điều cần đọc. Grounding lập chỉ mục là đang được xây dựng để help AI các hệ thống decide điều cần chẳng hạn.» — Madhavan, Risvik & Merchant, Microsoft AI (Có thể 2026). Nguồn
Academia — vì sao LLMs rerank well
- “Properly instructed LLMs can deliver competitive, even superior results to state-of-the-art supervised methods on popular IR benchmarks.” (bản dịch) «Properly instructed LLMs có thể deliver competitive, ngay cả superior kết quả để state-of-đó-art supervised các phương thức on popular IR benchmarks.» — Sun et al., “Is ChatGPT Good at Search?” (bản dịch) «Là ChatGPT Good tại Tìm kiếm?» (RankGPT), EMNLP 2023. arXiv
- LLMs có “a sense of pairwise relative comparisons, which is much simpler than requiring calibrated pointwise relevance estimation.” (bản dịch) «một hợp lý of pairwise relative comparisons, mà là nhiều simpler hơn requiring calibrated pointwise relevance estimation.» — Qin et al., Pairwise Xếp hạng Prompting. arXiv
- LLM rerankers exhibit “model-specific, non-uniform behavior that cannot be reduced to a single recognizable optimization strategy.” (bản dịch) «model-cụ thể, non-uniform behavior đó không thể là reduced để một single recognizable optimization strategy.» — “Diagnosing LLM Reranker Behavior Under Fixed Evidence Pools” (bản dịch) «Diagnosing LLM Reranker Behavior Dưới Fixed Evidence Pools» (2026). Đó caveat so với một single universal “vibe.” arXiv
Tự kiểm tra: Vibe xếp hạng
Các tài nguyên worth của bạn time
My related writing (Ahrefs)
- An Analysis of AI Overview Brand Visibility Factors (75K Brands Studied) — điều gì thực ra correlates với getting pulled vào AI Overviews: brand mentions (0,664), anchor text (0,527), lưu lượng tự nhiên (0,274). Đó visibility side of surviving đó reranking pipeline.
- Đó Great Decoupling — cách AI Overviews split impressions từ clicks; context cho vì sao surviving reranking ≠ getting đó visit.
- Websites Với Hơn Organic Tìm kiếm Traffic Nhận Mentioned Hơn trong AI Tìm kiếm — đó correlation giữa truyền thống visibility và AI mentions.
Foundational explainers (others)
- Pinecone — Rerankers và Hai-Stage Retrieval — đó clearest kỹ thuật explainer of đó pipeline, bi-encoder so với cross-encoder, và vì sao hai stages exist.
- ZeroEntropy — Nên Bạn Dùng an LLM as một Reranker? — benchmark numbers, đó cost reality of pointwise so với listwise, và một được khuyến nghị hybrid pipeline.
Academic papers
- Là ChatGPT Good tại Tìm kiếm? (RankGPT) — Sun et al., EMNLP 2023. Listwise reranking với sliding windows.
- LLMs là Effective Text Rankers với Pairwise Xếp hạng Prompting — Qin et al. Đó pairwise approach và của nó robustness.
- RankVicuna: Zero-Shot Listwise Reranking với Open-Nguồn LLMs — một 7B open model approaching GPT-3,5-cấp độ reranking.
- Diagnosing LLM Reranker Behavior Dưới Fixed Evidence Pools — đó model-specificity caveat.
- Passage Re-xếp hạng với BERT — Nogueira & Cho, 2019. Đó cross-encoder era này xây dựng on.
SEO-practitioner đọc
- Beyond RAG: Vì sao mỗi AI tìm kiếm nền tảng là hiện tại agentic — đó five-node agentic architecture và đó Google pairwise-patent đọc.
- Luca Tagliaferro — Google AI Pairwise Patent & Scorecard — practitioner analysis of điều gì wins pairwise comparisons.
My speaking
- Cách Tìm kiếm Hoạt động (SlideShare) — my walkthrough of crawling, kết xuất, lập chỉ mục, và xếp hạng. (Standing disclaimer: “This is my understanding of systems… not going to be 100% complete or accurate.” (bản dịch) «Này là my understanding of các hệ thống… không going để là 100% hoàn tất hoặc chính xác.»)
Từ khoảng đó ngành
- Search Engine Journal — Cách Researchers Reverse-Engineered LLMs cho một Xếp hạng Thử nghiệm — GPT-4o, Claude, Gemini, và Grok compared as rankers; covers điều gì các tín hiệu mỗi LLM implicitly weights.
- Search Engine Land — LinkedIn Cập nhật Feed Algorithm Với LLM Xếp hạng và Retrieval — một of đó đầu tiên major production confirmations of LLM reranking tại quy mô bên ngoài of web tìm kiếm; hữu ích thực tế precedent.
- ACL Anthology — PRP-Graph: Pairwise Xếp hạng Prompting với Graph Aggregation (ACL 2024) — extends đó PRP pairwise approach với graph-based aggregation; đó peer-reviewed follow-lên để đó original PRP paper.
- RankLLM: MỘT Python Toolkit cho LLM Reranking (SIGIR 2025) — open-nguồn toolkit cho experimenting với listwise và pairwise LLM reranking; đó practical research infrastructure underpinning hầu hết gần đây benchmarks.
- Google Blog — Tìm kiếm tại Google I/O 2026 — Gemini 3,5 Flash trong AI Chế độ; agentic on-đó-fly generation; chính thức cách diễn đạt of cách Google AI tìm kiếm pipeline là evolving.
Nhật ký thay đổi
Đã cập nhật 8 thg 8, 2026.
Tóm tắt biên tập và chi tiết thay đổi đã ghi nhận.Chi tiết thay đổi
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
Không thể so sánh đầy đủ — không có bản lưu trước đó cho lần sửa đổi này.
Đã cập nhật 19 thg 7, 2026.
Tóm tắt biên tập và chi tiết thay đổi đã ghi nhận.Chi tiết thay đổi
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
Không thể so sánh đầy đủ — không có bản lưu trước đó cho lần sửa đổi này.