natural sentences or answer questions, much like predicting the word “sunny” for the input “The weather today is...”. This technology serves as the foundation for generative AI services such as ChatGPT. Foundation Model “Foundation Model” is a term used for large-scale AI models that have been pre-trained on massive datasets and can be adapted for a wide variety of tasks. In the field of generative AI, it refers to the base model before it is fine-tuned for specific tasks such as text generation, image creation, or speech synthesis. For instance, GPT4 and Stable Diffusion are popular foundation models for language and image generation, respectively, and generative AI services like ChatGPT operate based on these models. Multimodal Multimodal(ity) refers to the capability of simultaneously processing or generating data in different forms, such as text, images, audio, and video. Examples of systems that adopt a multimodal approach include AI that generates images based on text descriptions (like Midjourney) or describe the content of an image in text form, or systems that understand voice commands to recommend video content. Modality is a term from semiotics referring to the forms of communication like writing, images, or music; in the AI context, it can be understood as synonymous with data format. Multimodal models perform tasks that cannot be accomplished by models 16 17

Select target paragraph3