← Curriculum map
L18 · Phase 5 · 20 min

Multimodal AI

How can AI Tool look at my photo of the boiler room and discuss what's actually in it?

Multimodal Explorer — boiler-room photo

A stand-in for a real photo, divided into patches the way Attention divides a sentence into tokens. Click a question to see which patch the model should attend to.

Prototype note: this is an illustrated stand-in, not a real uploaded photo run through a real vision model — the patch layout is hand-drawn to make the tokens-for-images idea concrete, not extracted from actual image processing.

Depth ladder

A photo of the boiler room gets broken into small square pieces the model can relate to words, in much the same way your message gets broken into text tokens — that's what lets the model 'read' an image at all.

Knowledge check

Given a condo photo and a question about it, identify which region of the image the answer should depend on.