GPT-4o System Card
OpenAI Thanks: Please cite this work as “OpenAI (2024)". Full authorship contribution statements appear at the end of the document.
Abstract
GPT-4o[1] is an autoregressive omni model, which accepts as input any combination of text, audio, image, and video and generates any combination of text, audio, and image outputs. It’s trained end-to-end across text, vision, and audio, meaning that all inputs and outputs are processed by the same neural network.
原文 arXiv:2410.21276;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2410.21276v1