For a long time a model could only read text. Everything else, a photo of an invoice, a screenshot, a call recording, had to be converted by hand before it could be used. That intermediate step has just disappeared, and for many small businesses it is the most concrete change of the past two years.
Definition
A multimodal model handles several forms of content within the same conversation: text, image, sound, sometimes video. You can show it a photo and ask a written question about it.
The nuance that matters: this is not several tools chained together, where an image would first be described in text before being analysed. The model processes the image directly, which lets it grasp layout, formatting and details a description would lose.
The difference is very visible on documents. A multimodal model understands that a value belongs to a column of a table, or that an amount is the total rather than a line item. A plain-text transcription would have lost that structure.
The uses that change life in a small business
| Situation | What it replaces |
|---|---|
| Photo of an invoice or receipt | Manual entry of the amounts |
| Screenshot of an error | An approximate description of the problem |
| Recording of a meeting | Note taking and minutes |
| Photo of a handwritten document | Retyping by hand |
| A layout sketched on paper | The specification to write up |
What these share: they are all conversion tasks, not creation. That is where reliability is highest, because the material is given and there is nothing to invent.
What to watch
Images consume a lot. An image is counted in Token, often the equivalent of several hundred words. An automation handling photos costs noticeably more than a text automation, which has to be anticipated in the budget.
Reading remains fallible. A badly printed figure, cramped handwriting, a crooked photo: reading errors happen and arrive with the same confidence as everything else. On amounts, verification is required.
Confidentiality works differently. A photo often contains more than what you meant to show: the rest of the desk, another document, a lit screen. Look at what surrounds the subject before sending.
Frequently asked questions
Does it replace text recognition software?
For occasional and varied use, largely. For mass processing of one document type, a specialised tool is often faster, cheaper and more consistent. Multimodal shines on diversity, the dedicated tool on repetition.
Can a model also produce images?
Reading and producing are two distinct capabilities. Many models can analyse an image without being able to create one, and the reverse exists too. Check which direction you need before choosing.
Is video handled like an image?
Generally, video is cut into still frames, accompanied by the soundtrack. That works well for understanding content, less well for fast movement or a detail falling between two frames.
Where do you start to benefit from it?
With the data entry task that bores you most: it is almost always the one that converts best. Our Claude Cowork course starts from these concrete cases, invoices and meeting notes first.