Multimodal Tìm kiếm

Cách tìm kiếm hoạt động trên text, images, video, và audio — Google Lens, Circle để Tìm kiếm, multimodal AI Chế độ, và MUM — và điều gì các công cụ tìm kiếm có thực ra confirmed về optimizing cho điều này.

Xuất bản lần đầu: 3 thg 7, 2026 · Cập nhật lần cuối: 8 thg 8, 2026 · Advanced
Ngôn ngữ

Multimodal tìm kiếm có nghĩa là querying và retrieving trên hơn một modality — text, images, video, và audio together — thay vì typing words và getting links. đây là đó shift behind Google Lens (point của bạn camera), Circle để Tìm kiếm (circle điều gì đó on screen), voice tìm kiếm, và multimodal AI Chế độ, nơi một photo plus một follow-lên câu hỏi becomes một query. Đó kỹ thuật enabler là shared embeddings: images, text, và audio nhận mapped vào đó giống nhau vector space so they có thể là khớp by meaning, đó giống nhau mechanism đó powers semantic tìm kiếm extended past text. Google MUM (2021) đã là đó công khai marker cho này; Gemini-era AI Chế độ pushes điều này further. điều gì là officially confirmed là mostly về understanding đó query trên modalities — sweeping 'optimize cho multimodal xếp hạng' claims là ngành theory, không confirmed mechanism. Đó durable playbook là unglamorous: descriptive alt text và filenames, real transcripts và captions, sạch dữ liệu có cấu trúc, và nội dung đó các câu trả lời đó câu hỏi một visual hoặc spoken query implies.

TL;DR — Multimodal tìm kiếm là querying và retrieving trên text, image, video, và audio trong một single interaction. Đó enabler là một shared embedding space: khác nhau modalities nhận encoded vào đó giống nhau vector space so they có thể là khớp by meaning — đó giống nhau retrieval ý tưởng as semantic tìm kiếm, generalized past text. Google công khai arc chạy từ voice và image tìm kiếm để MUM (2021, multimodal, “1,000 times more powerful than BERT,” (bản dịch) «1 000 times hơn powerful hơn BERT,» dùng cho hẹp cao-complexity cases) để Google Lens, Circle để Tìm kiếm, và Gemini-powered AI Chế độ, nơi một photo plus một follow-lên là một query đó fans out vào một synthesized câu trả lời. điều gì là officially confirmed là mostly về understanding đó query trên modalities; claims đó Google ranks của bạn riêng lẻ images hoặc video frames as một distinct “multimodal signal” (bản dịch) «multimodal tín hiệu» là ngành theory. Đó durable playbook: descriptive alt text/filenames, real transcripts/captions, sạch dữ liệu có cấu trúc, và nội dung đó các câu trả lời đó implied câu hỏi behind một visual hoặc spoken query.

Điều gì “multimodal” thực ra có nghĩa là

Shared representation research cho thấy một approach để cross-modal retrieval, không một universal production architecture. Evidence for this claim CLIP learns joint image-and-text representations from paired internet data and supports zero-shot image classification. Scope: CLIP research results; multimodal search products may use different models, training data, and ranking pipelines. Confidence: high · Verified: Radford et al.: Learning Transferable Visual Models Tìm kiếm-sản phẩm các mô tả document features nhưng không hoàn tất xếp hạng mechanics. Evidence for this claim Google Multisearch lets users combine an image with text to refine a search. Scope: Google's documented consumer search feature; it does not define all multimodal retrieval systems. Confidence: high · Verified: Google: Multisearch

MỘT modality là một channel of information — được viết text, an image, spoken audio, video. MỘT unimodal hệ thống xử lý một; một multimodal hệ thống takes hơn một as input, hoặc produces hơn một as output, và — crucially — reasons về them jointly thay vì bolting tách biệt pipelines together.

Multimodal tìm kiếm là rộng hơn “image search,” (bản dịch) «image tìm kiếm,» và đó khác biệt matters:

  • Image tìm kiếm tìm thấy images (bạn muốn pictures).
  • Multimodal tìm kiếm dùng an image (hoặc voice, hoặc video) as part of một query đó có thể trả về bất kỳ kind of kết quả — một shopping listing, một cách-để, một definition, một map pin. Đó camera là an input device, không đó điều bạn là looking cho.

So “point your camera at a broken part and ask how to fix it” (bản dịch) «point của bạn camera tại một hỏng part và ask cách sửa điều này» là multimodal tìm kiếm; “find me photos of that part” (bản dịch) «tìm me photos of đó part» là image tìm kiếm. Google Lens làm cả hai, mà là part of vì sao đó line blurs.

Điều này cũng helps để giữ three tách biệt các câu hỏi apart, vì một single sản phẩm mô tả có thể âm thầm blur them: điều gì modality có thể đó query accept (text chỉ, hoặc text plus image, audio, video)? Điều gì modality là đó retrieved candidate nội dung trong (một text trang, an image, một video)? Và điều gì modality là đó generated output (một list of links, an image grid, một synthesized text câu trả lời)? MỘT feature đó accepts an image input không tự động retrieve image kết quả hoặc produce an image output — một photo of một hỏng part có thể retrieve một text trang và generate một text câu trả lời. Confirming một of những three không xác nhận đó other hai; kiểm tra mỗi một so với điều gì đó provider có thực ra được ghi lại cho đó cụ thể feature.

Đó kỹ thuật enabler: một shared space

Multimodal retrieval works when different input types can be compared by meaning in one shared space. Nguồn: /ai-search/how-search-works/embeddings/

Text queries, image queries, and audio or video queries are encoded into a shared semantic embedding space. Meaningfully similar items land near one another, allowing a nearest-neighbor search to retrieve relevant content across input types. The diagram is a conceptual model, not a claim about a specific Google ranking formula.

© Patrick Stox LLC · CC BY 4.0 ·

Đó reason một picture và một phrase có thể là compared tại all là shared representation. Encoder models map text, images, audio, và video vào đó giống nhau cao-dimensional embedding space, nơi semantic similarity becomes geometric closeness. MỘT photo of một plant và đó string “how do I care for this plant” (bản dịch) «cách làm I care cho này plant» land near mỗi other, so một nearest-neighbor lookup có thể connect them.

Này là chính xác đó machinery behind semantic tìm kiếm và đó retrieval leg of RAG, generalized beyond text, và điều này breaks vào distinct stages — mỗi một tách biệt khả năng, so evidence đó một hệ thống là good tại một stage không xác nhận đây là good tại đó others:

  1. Input understanding. Đó query (image + text, chẳng hạn, hoặc một spoken câu hỏi) là parsed vào whatever representation đó tiếp theo stage cần.
  2. Embedding. Đó representation, và mỗi candidate piece of nội dung, là encoded vào đó shared vector space described trên.
  3. Retrieval. Vector tìm kiếm tìm thấy đó closest matches by geometric distance.
  4. Reranking. MỘT tách biệt stage reorders đó retrieved candidates dùng các tín hiệu beyond thô embedding distance.
  5. Generation. Trong AI surfaces, đó top candidates là grounded vào một synthesized câu trả lời.

Nếu bạn understand embeddings, chunking, vector tìm kiếm, và grounding, bạn đã understand multimodal tìm kiếm plumbing — đó chỉ new part là đó encoders speak hơn một modality.

Named research các hệ thống, không một confirmed production blueprint. Đó joint-embedding ý tưởng không hypothetical — OpenAI’s CLIP và Google ALIGN cả hai trained image và text encoders vào một shared space dùng image-text pairs, và Meta ImageBind extended đó approach để six modalities (image, text, audio, depth, thermal, và motion) trong một single space. Những là named, published research architectures với của họ own datasets và evaluations — hữu ích cho understanding cách một shared space có thể là được xây dựng, không proof đó bất kỳ cụ thể commercial công cụ tìm kiếm production hệ thống dùng đó chính xác design. MỘT provider output behavior alone không reveal liệu điều này dùng một shared embedding space, một khác nhau fusion phương thức, hoặc cách điều này weights mỗi modality — đó là internal để một hệ thống không ai bên ngoài đó provider có thể inspect.

Google công khai arc

Google confirmed steps toward multimodal tìm kiếm, khoảng trong order:

  • Voice và image tìm kiếm (2010s). Spoken các truy vấn và reverse-image lookups đã là đó đầu tiên mainstream non-text inputs.
  • MUM — Multitask Unified Model (2021). Google công khai marker cho đó multimodal shift. Google described MUM as multimodal — able để “understand information across text and images” (bản dịch) «understand information trên text và images» — và “1,000 times more powerful than BERT,” (bản dịch) «1 000 times hơn powerful hơn BERT,» trained trên 75 languages, và able để cả hai understand generate language. Đó quan trọng caveat, straight từ đó semantic-tìm kiếm history: MUM đã là không bao giờ Google chung xếp hạng engine. Điều này đã là applied để hẹp, cao-complexity cases (phức tạp multi-step các câu hỏi, some featured snippets, shopping), và điều này đã không “replace” BERT.
  • Google Lens. Camera-as-query tại quy mô — identify objects, translate text trong an image, shop điều gì bạn see, solve một vấn đề bạn photograph.
  • Circle để Tìm kiếm. Circle, highlight, hoặc tap bất cứ điều gì on của bạn Android screen để tìm kiếm điều này trong place, không có switching apps.
  • AI Chế độ (Gemini-era). Đó hiện tại frontier: một multimodal query (an image plus một spoken hoặc typed follow-lên) becomes một single interaction đó fans out vào nhiều sub-các truy vấn và trả về một synthesized, cited câu trả lời. Này là nơi multimodal input đáp ứng đó agentic, RAG-style retrieval described trong này cluster.

Microsoft có của nó own line — Bing Visual Tìm kiếmCopilot Vision — đó làm đó analogous điều on Bing side, so multimodal tìm kiếm không một Google-chỉ phenomenon.

Nơi điều này có thể fail trước retrieval ngay cả bắt đầu

Mỗi modality other hơn sạch text có để truyền qua một conversion step trước đó shared-space matching described trên có thể happen, và đó step có thể introduce các lỗi đó retrieval stage không bao giờ sees coming. Provider kỹ thuật tài liệu on multimodal models mô tả chính xác này class of limitation:

  • OCR các lỗi — text embedded trong an image (một sign, một screenshot, một label) có thể là misread, especially tại thấp resolution hoặc trong an unfamiliar font.
  • Speech recognition các lỗi — một mistranscribed word trong audio hoặc video propagates vào whatever nhận embedded và retrieved.
  • Frame sampling — video không understood frame-by-frame; các hệ thống sample một subset of frames, so nội dung đó chỉ xuất hiện giữa sampled frames có thể là missed hoàn toàn.
  • Cropping và resolution — một hệ thống hoạt động từ một thấp-resolution hoặc tightly cropped image có ít hơn để hoạt động với hơn đó original scene.
  • Language — độ chính xác trong một language không establish độ chính xác trong một sản phẩm khác; OCR và speech recognition performance cả hai vary by language.
  • Bị thiếu context — an image hoặc clip stripped of của nó xung quanh trang context (caption, alt text, transcript) cho đó hệ thống ít hơn để anchor understanding để.

None of này là unique để bất kỳ một provider, và đây là một reason để treat “the model can see/hear X” (bản dịch) «đó model có thể see/hear X» claims cautiously — một khả năng đó hoạt động on một sạch, well-lit, cao-resolution kiểm thử ví dụ không establish điều này hoạt động on đó messier thực tế version of đó giống nhau input.

điều gì là confirmed so với. điều gì là theory

Này là nơi careful writing matters, vì đó topic attracts overreach.

Reasonably confirmed:

  • Các công cụ tìm kiếm accept và understand multimodal các truy vấn — image, voice, và screen-circling inputs là real, shipping các sản phẩm.
  • Đó retrieval behind AI các câu trả lời là meaning-based (embeddings / semantic retrieval) và, trong AI Chế độ, grounded trong đó cốt lõi Tìm kiếm chỉ mục — có không tách biệt “AI chỉ mục.”
  • Google explicitly says của nó AI features rely on cốt lõi Tìm kiếm xếp hạng và quality các hệ thống, so đó crawl → chỉ mục → retrieve chain vẫn governs liệu nội dung của bạn có thể cho thấy lên tại all.

Ngành theory — flag điều này as such:

  • Đó Google ranks của bạn riêng lẻ images hoặc video frames as một distinct “multimodal ranking signal” (bản dịch) «multimodal tín hiệu xếp hạng» bạn có thể optimize trong isolation. Google có không published một “multimodal ranking factor.” (bản dịch) «multimodal xếp hạng factor.» Cách deeply trang-cấp độ visual/audio nội dung là understood frame-by-frame tại web quy mô là chỉ partly được ghi lại.
  • Precise claims về cách nhiều weight một multimodal query cho để đó image so với. đó accompanying text. Treat cụ thể ratios as speculation.

Đó safe reading: optimize cho đang understood và retrievable trên modalities, không cho một mechanism Google hasn’t confirmed.

Điều gì này có nghĩa là cho SEO

Strip đó hype và đó playbook là concrete và familiar:

  • Làm images legible để machines. Descriptive alt text, có ý nghĩa filenames, và (nơi relevant) image dữ liệu có cấu trúc. Alt text là cách một tìm kiếm engine knows điều gì một picture depicts không có “seeing” điều này — và điều này vẫn helps ngay cả as visual understanding improves.
  • Cho audio và video real text. Transcripts, captions, và clear on-trang context turn spoken nội dung vào điều gì đó retrievable và quotable. MỘT video với không transcript là far harder để ground vào an câu trả lời.
  • Dữ liệu có cấu trúc nơi điều này fits. Product, ImageObject, VideoObject, và Recipe-loại markup cho machines unambiguous facts để attach để một visual hoặc spoken query.
  • Câu trả lời đó implied câu hỏi. MỘT visual query carries an intent — “what is this,” (bản dịch) «điều gì là này,» “where do I buy it,” (bản dịch) «nơi làm I buy điều này,» “how do I fix it.” (bản dịch) «cách làm I cách sửa điều này.» Nội dung đó các câu trả lời đó intent trực tiếp, trong một self-contained passage, là điều gì nhận retrieved. Này là đó giống nhau passage-cấp độ clarity đó passage xếp hạng và RAG đã reward.
  • Đó prerequisites không thay đổi. Multimodal các câu trả lời vẫn retrieve từ đó cốt lõi chỉ mục, so đang crawlable và được lập chỉ mục xuất hiện đầu tiên — đó giống nhau chain covered trong crawling, grounding, và RAG.

None of này là một bảo đảm. Google own image và video tài liệu frames alt text, transcripts, và dữ liệu có cấu trúc as eligibility và understanding aids — điều đó help một công cụ tìm kiếm discover, understand, và correctly attach của bạn nội dung để một query — không as inputs để một được ghi lại xếp hạng formula. Đang làm all of đó trên làm nội dung của bạn retrievable và citable; điều này không bảo đảm retrieval, citation, xếp hạng position, hoặc đó bất kỳ generated câu trả lời mô tả điều này accurately.

Đó một-sentence version: multimodal tìm kiếm widened đó front door, nhưng điều này đã không thay đổi điều gì là behind điều này. Là findable, là clear về điều gì của bạn images và media depict, và câu trả lời đó câu hỏi đó query implies.

Add an expert note

Pin an expert quote

New person? Create their unclaimed profile at /admin/experts/ → Pin a quote first.