As Google continues to cycle through its products to give them AI upgrades and infusions, the latest on the list is Google Lens. The visual search tool now boasts more intelligence to not only identify items in a photo or live camera feed but also let users go deeper in asking questions about said items.
This is Google’s latest form of multisearch. As its name implies, it lets users search using several inputs and formats. For example, after a Google Lens search to identify a physical object using your camera, you can refine or filter results with text, (e.g., “the same jacket in green.”), thus greater optionality.
As background for those unfamiliar, Google Lens is the search giant’s visual search play. Rather than typing (or speaking) text, this visual modality lets you simply point your camera at objects to identify or contextualize them. You can also do this with photos on your device or those encountered on the web.
This isn’t a silver bullet, nor will it fully replace deep-rooted habits around traditional search, but it does offer a more intuitive search front end in some contexts. For example, fitting use cases include items in nature (think: plant species), shopping (think: fashion items), and local discovery (think: storefronts).
Input/Output
Back to the most recent updates, users can now go deeper into visual searches as noted. This was possible before using multisearch, which returned refined visual results. The difference with the newest update is that generative AI joins the party to return additional text-based insights as well.
So in summary, using the same example above, you can take a picture of a jacket with Google Lens to see images of the same or visuall-similar items. You can then refine those visual results (again “the same jacket in green”) to get updated results. Now you also get text telling you that it’s part of H&M’s fall line.
Applying that to a local search, someone in a new neighborhood can use Google Lens to identify a storefront. They can then ask several qualifying questions about business attributes. Is the restaurant pet-friendly? Does it accommodate large groups for a birthday party? What’s its Yelp rating?
All this data flows from Google Business Profiles. Surfacing it in the Google Lens clickstream is conceptually similar to what Google has been doing for years with the knowledge panel in core search results. More broadly, multisearch is an evolution of what Google used to call “universal search.”
Zeroing in on the latter, universal search was the early 2010s evolutionary step when Google results included a mix of text, video, images, and shopping. The difference now is that the search input formats are varied in addition to the output formats. That’s multisearch in a nutshell… now with more visual AI.

Meet the New AI…
Backing up, one thing that jumps out here is that Google Lens is inherently AI-oriented. On the back end, its core functionality – recognizing and contextualizing physical-world objects – is accomplished through Google’s knowledge graph. In this case, 20 years of Google Images represent one giant training set.
So in essence, Google Lens’ latest AI infusions add AI on top of AI. More accurately, it upgrades the product’s existing AI with some of the newer flavors of AI that have rapidly advanced in the past year. These include large visual models (like large language models but image-based) that fuel generative AI.
Buried in all of this is an ongoing lesson of the recent AI hype cycle: AI is not new. Though the technology has inflected in interest and investment – and in capability – the message often sent in generalist tech media and beyond is that it’s some new scary thing. In reality, AI is as mundane as spell check.
That said, recent AI inflections are real. The term ‘hype cycle’ above is perhaps unfair because AI isn’t as hollow as recent hype cycles (we’re looking at you, metaverse). The technology is here today, and broadly applicable. So expect more infusions and upgrades for Google and the rest of the tech universe.


