Signal & Noise
Menu

Multimodal

Models that accept more than text — images, audio, video — in the same conversation. Image inputs are tokenized too, and cost real money per image.

← Back to the full glossary