AIWiser LogoHow To Talk To AIChapter 2 of 4
Course outline

What is Multimodal Prompting?

3 min read

Most people only ever type text into an AI tool. Multimodal prompting goes further by combining text, images, and audio to communicate more clearly and get more accurate results. Modalities are simply different input formats, for example text, image, audio, or video. Multimodal prompting means using more than one of these together in a single prompt. For example, uploading an image alongside a written instruction, or describing a sound while referencing an existing audio file.

Why should you use it?

Because some things are easier to show than explain. Describing a complex graph in words is far harder than just sharing it. Multimodal prompting bridges that gap by making your communication with AI faster, clearer, and more precise.

2 key uses

  • Reference Past Work: Upload existing designs, documents, or outputs and ask the AI to build on them. It picks up your style, colours, and preferences automatically.
  • Image to Text: Upload an image, get a detailed description. Use it to generate similar visuals or show the AI exactly what you mean.

The golden rule

Don't just upload a file. You need to explain what it is, why it's there, and what you need.

Here is a bad and good example of using multimodal prompting. Let's say you want to create more logo designs and you have a reference photo.

  • Bad example: Upload a logo and then write "Make me something like this logo."
  • Good example: A better way is to upload a logo then write a written prompt: "This is my current brand logo. Create three new variations that maintain the same minimalist style and navy blue and gold colour palette, but with a more modern geometric approach. Each version should work at small sizes for social media icons."

The second prompt gives context, explains the purpose, and gives the AI clear direction on how to use the visual reference.

When does it work best?

Multimodal prompting is most powerful when you're transferring information between formats:

  • Image to text: describing or analysing a photo
  • Audio to text: transcription and summarisation
  • Text to image: generating visuals from a description
  • Image to image: applying a style or making variations

Best practices

Start with one or two modalities, not five at once. Be explicit about what each input is for. Provide context so the AI understands how the pieces relate. Test, refine, and iterate based on your results. And always match your modalities to the task at hand.

Common mistakes to avoid

  1. Uploading files without explaining what they are or why they're included.
  2. Using too many modalities at once and confusing the AI.
  3. Expecting the AI to read your mind. Be explicit about connections between inputs.
  4. Not testing different combinations to see what works best.

Start multimodal prompting

Multimodal prompting opens up new possibilities for working with AI. By combining text, images, audio, and other formats, you can communicate more effectively and get better results. Start experimenting with different modalities, keep it simple, and watch your AI outputs improve!