Part 2: Inside the mechanics of multimodal AI 

What does AI see when it looks at your product?  

What does AI see when it looks at your product?

Part 2: Inside the mechanics of multimodal AI

 

In Part 1 we discussed the fact that consumers are increasingly using AI to find, select and compare products for them. AI Blog Header sectionWe explored how both product data and images are becoming signals that AI can interpret when identifying and understanding products, rather than just promotional materials which appeal to the human eye.

But what does that reading involve? If AI cannot see as we can, feel a physical product or have a brand preference that influences a decision, how does it interpret and understand your products?

This is where it gets that little bit more technical, so let’s dive in.

 

Meet multimodal AI

 

An increasing number of consumer-facing AI tools, shopping assistants, retailer search bars and chatbots, use multimodal AI. Unlike older, single-input models, it interprets several data types simultaneously, including text, images, and audio, and then combines all into a single understanding. Before a decision is then made, a vision model examines the photos and a language model reads the pack copy and description, and then both compare notes to establish a verdict.

However, exactly how this happens varies between different models and platforms.

 

How does a vision model read images

 

Unlike humans, vision models don't ‘see’ a jar of coffee in the same way we do. Computer vision models can analyse visual characteristics within an image, including patterns, colours, forms, edges, textures and proportions.

 

Depending on the model and task, this can enable it to:

 

    • Identify a product or product category

    • Identify visual similarities between products

    • Use image recognition to extract visible on-pack text

    • Associate visual characteristics with concepts

    • Support product classification and visual searchHIPP_Organic_Case_Study-1

We are already seeing this used in shopping.

For example, Google Lens allows consumers to photograph an object and find visually similar products.

Another example is OpenAI's CLIP, which learns links between text and images. This helps demonstrate how an AI system can connect what appears within an image with relevant language and concepts, rather than relying just on what's in a written description.

However, CLIP is not necessarily the technology behind any particular shopping assistant. Instead, it's a useful example of how AI has the ability to connect what appears within an image with language and meaning.

 

Why does that matter for your product imagery?

 

It means imagery is increasingly capable of functioning as information, as well as creative content.

But that doesn't mean cleaner imagery automatically leads to a higher AI ranking or recommendation. Different platforms use different models, data sources and ranking systems, and there is no evidence for a simple formula where:

Better image = AI recommends your product.

What we do know is that the quality of information supplied to AI matters. Poor-quality, inaccurate or conflicting inputs can reduce the reliability of multimodal systems.

For brands, that creates another reason to ensure the product information they put into the digital ecosystem is as clear, accurate and reliable as possible.

 

Key practical takeaways

 

“Good” product imagery is just not good enough anymore, it also needs to be:

 

  • Consistent across every channel and retailer (basically everywhere your products can be seen)

  • Accurate to the current pack, with no missing or outdated information.

  • High quality and clear, so that colours, shapes, textures and on-pack text can extract cleanly

  • Free of visual noise that could impact the reading and lead to being misread as product features

These qualities also tend to lend themselves to a better shopping experience for human shoppers. Clear, accurate and consistent product information benefits people trying to understand what they're buying, while also providing machines with more reliable information to interpret.

 

Where the challenge really sits

 

For an FMCG brand managing a handful of products, keeping imagery accurate might be relatively straightforward. Now scale that across hundreds or thousands of SKUs, multiple retailers, marketplaces, regional variants and regular pack changes.

Suddenly, producing the image is only one part of the challenge.

You also need to know which version is current, whether the latest artwork has been used, whether the correct assets have been approved and whether they're being delivered to the right places.

AI doesn’t create these challenges, but as machines become more advanced in interpreting the information brands put onto the digital shelf, those qualities outlined above become even more important.

For FMCG brands, that means having the right processes and technology around the creation, approval, management and distribution of product assets at scale.

 

The bottom line

women holding a ipad looking at english provender products

This doesn’t mean brands need to create imagery for an algorithm. It means your product imagery is increasingly being interpreted by machines as well as people.

There's no guarantee that cleaner or more consistent imagery will make an AI systems recommend your product over a competitor.

But giving both people and machines clear, accurate and reliable information puts your product in the best position to be correctly understood.

And as AI becomes a bigger part of how consumers discover products, that's something FMCG brands can't afford to overlook.

 

 

Reference list

API4AI (2025). 5 computer vision tactics to boost e-commerce visibility. [online] Medium. Available at: https://medium.com/@api4ai/5-computer-vision-tactics-to-boost-e-commerce-visibility-9220b9c229a5 

Google. Google Lens visual search and shopping resources.

Misikir Tashu, T., Fattouh, S., Kiss, P. and Horváth, T. (2022). Multimodal e-commerce product classification using hierarchical fusion. [online] Available at: https://arxiv.org/pdf/2207.03305 

Nguyen, D.A. (2026). Multimodal AI Examples: How It Works, Real-World Applications, and Future Trends. [online] SmartDev. Available at: https://smartdev.com/multimodal-ai-examples-how-it-works-real-world-applications-and-future-trends/#3_How_Multimodal_AI_Works 

OpenAI (2021). clip: connecting text and images. [online] Openai.com. Available at: https://openai.com/index/clip/ 

 

Written By Gaby Oldham