Multimodal: an AI that also reads images and sound

A multimodal model handles several types of content at once: text, image, sound, video, in the same conversation.
3 min read
Believemy logo

For a long time a model could only read text. Everything else, a photo of an invoice, a screenshot, a call recording, had to be converted by hand before it could be used. That intermediate step has just disappeared, and for many small businesses it is the most concrete change of the past two years.


Definition

A multimodal model handles several forms of content within the same conversation: text, image, sound, sometimes video. You can show it a photo and ask a written question about it.

The nuance that matters: this is not several tools chained together, where an image would first be described in text before being analysed. The model processes the image directly, which lets it grasp layout, formatting and details a description would lose.

Good to know

The difference is very visible on documents. A multimodal model understands that a value belongs to a column of a table, or that an amount is the total rather than a line item. A plain-text transcription would have lost that structure.


The uses that change life in a small business

SituationWhat it replaces
Photo of an invoice or receiptManual entry of the amounts
Screenshot of an errorAn approximate description of the problem
Recording of a meetingNote taking and minutes
Photo of a handwritten documentRetyping by hand
A layout sketched on paperThe specification to write up

What these share: they are all conversion tasks, not creation. That is where reliability is highest, because the material is given and there is nothing to invent.


What to watch

Images consume a lot. An image is counted in Token, often the equivalent of several hundred words. An automation handling photos costs noticeably more than a text automation, which has to be anticipated in the budget.

Reading remains fallible. A badly printed figure, cramped handwriting, a crooked photo: reading errors happen and arrive with the same confidence as everything else. On amounts, verification is required.

Confidentiality works differently. A photo often contains more than what you meant to show: the rest of the desk, another document, a lit screen. Look at what surrounds the subject before sending.


Frequently asked questions

Question

Does it replace text recognition software?

For occasional and varied use, largely. For mass processing of one document type, a specialised tool is often faster, cheaper and more consistent. Multimodal shines on diversity, the dedicated tool on repetition.


Question

Can a model also produce images?

Reading and producing are two distinct capabilities. Many models can analyse an image without being able to create one, and the reverse exists too. Check which direction you need before choosing.


Question

Is video handled like an image?

Generally, video is cut into still frames, accompanied by the soundtrack. That works well for understanding content, less well for fast movement or a detail falling between two frames.


Question

Where do you start to benefit from it?

With the data entry task that bores you most: it is almost always the one that converts best. Our Claude Cowork course starts from these concrete cases, invoices and meeting notes first.

Related terms

Discover our aI and automation glossary

The vocabulary of artificial intelligence and automation, explained for people who want to use it in their business, not for people who build the models.

Share this article

Want to help us? Share this article on your networks or even better: on your site, in an article or in your newsletter.