natural sentences or answer questions, much like predicting the
word “sunny” for the input “The weather today is...”. This technology
serves as the foundation for generative AI services such as
ChatGPT.
Foundation Model
“Foundation Model” is a term used for large-scale AI models that
have been pre-trained on massive datasets and can be adapted for
a wide variety of tasks. In the field of generative AI, it refers to the
base model before it is fine-tuned for specific tasks such as text
generation, image creation, or speech synthesis. For instance, GPT4 and Stable Diffusion are popular foundation models for language
and image generation, respectively, and generative AI services like
ChatGPT operate based on these models.
Multimodal
Multimodal(ity) refers to the capability of simultaneously processing
or generating data in different forms, such as text, images, audio,
and video. Examples of systems that adopt a multimodal approach
include AI that generates images based on text descriptions (like
Midjourney) or describe the content of an image in text form, or
systems that understand voice commands to recommend video
content. Modality is a term from semiotics referring to the forms
of communication like writing, images, or music; in the AI context,
it can be understood as synonymous with data format. Multimodal
models perform tasks that cannot be accomplished by models
16
17