GPT-4o: Difference between revisions
Tecno.kids87 (talk | contribs) →Removal with GPT-5: I changed the tone to neutral Tags: Reverted references removed Mobile edit Mobile web edit |
→Capabilities: Revised the Capabilities section for clarity and compliance with the Manual of Style; removed unnecessary language parameters, consolidated overlapping fine-tuning sections, corrected promotional wording, and distinguished preprints from established findings. Tags: Reverted Mobile edit Mobile web edit |
||
| Line 77: | Line 77: | ||
{{See also|GPT Image}} |
{{See also|GPT Image}} |
||
GPT-4o is a multimodal model |
GPT-4o is a multimodal model capable of processing text, images, and audio. At its launch, it was described as integrating text, voice, and visual interaction within ChatGPT, rather than presenting them solely as separate features.<ref>{{Cite web |last=Wiggers |first=Kyle |title=OpenAI debuts GPT-4o 'omni' model now powering ChatGPT |url=https://techcrunch.com/2024/05/13/openais-newest-model-is-gpt-4o/ |website=TechCrunch |date=2024-05-13 |access-date=2026-06-30}}</ref><ref>{{Cite web |last=Fried |first=Ina |title=OpenAI debuts new model with enhanced real-time voice abilities |url=https://www.axios.com/2024/05/13/openai-google-chatgpt-ai |website=Axios |date=2024-05-13 |access-date=2026-06-30}}</ref> |
||
=== Text === |
=== Text === |
||
GPT-4o was introduced as a successor to GPT-4 Turbo in OpenAI's GPT-4 model family. Contemporary reports highlighted its response speed, lower operating cost, multilingual capabilities, and broader availability in ChatGPT compared with previous GPT-4-based offerings.<ref>{{Cite web |title=New GPT-4o AI model is faster and free for all users, OpenAI announces |url=https://www.theguardian.com/technology/article/2024/may/13/openai-new-chatgpt-free |website=The Guardian |date=2024-05-13 |access-date=2026-06-30}}</ref><ref>{{Cite web |title=OpenAI launches new 'flagship' GPT-4o model, touting improved speed and performance |url=https://www.marketwatch.com/story/openai-launches-new-flagship-gpt-4o-model-touting-improved-speed-and-performance-f7f40369 |website=MarketWatch |date=2024-05-13 |access-date=2026-06-30}}</ref> |
|||
A 2024 |
A 2024 evaluation assessed GPT-4o on language, vision, speech, and multimodal tasks, treating it as a general-purpose multimodal large language model rather than solely as a text-based conversational system.<ref>{{Cite arXiv |last=Shahriar |first=Sakib |last2=Lund |first2=Brady |last3=Mannuru |first3=Nishith Reddy |last4=Arshad |first4=Muhammad Arbab |last5=Hayawi |first5=Kadhim |last6=Bevara |first6=Ravi Varma Kumar |last7=Mannuru |first7=Aashrith |last8=Batool |first8=Laiba |title=Putting GPT-4o to the Sword: A Comprehensive Evaluation of Language, Vision, Speech, and Multimodal Proficiency |eprint=2407.09519 |class=cs.AI |date=2024}}</ref> |
||
=== Image === |
=== Image === |
||
GPT-4o supports image input |
GPT-4o supports image input, allowing users to submit photographs, charts, screenshots, and other visual materials for analysis. During its May 2024 launch demonstrations, the model was shown interpreting visual information, including mathematical problems and content displayed on a computer screen.<ref>{{Cite web |title=OpenAI launches GPT-4o, improving ChatGPT's text, visual and audio capabilities |url=https://apnews.com/article/071dd86594bc310bac07b37cf5e9bafc |website=Associated Press |date=2024-05-13 |access-date=2026-06-30}}</ref><ref>{{Cite web |title=OpenAI's big event: CTO Mira Murati announces GPT-4o, which gives ChatGPT a better voice and eyes |url=https://www.businessinsider.com/openai-event-live-updates-sam-altman-announcement-chatgpt-news-2024-5 |website=Business Insider |date=2024-05-13 |access-date=2026-06-30}}</ref> |
||
In March 2025, image generation |
In March 2025, OpenAI introduced image generation and editing through GPT-4o in ChatGPT. The feature was made available across several ChatGPT subscription tiers and was positioned as an alternative to the service's earlier DALL-E-based image-generation system.<ref>{{Cite web |last=Wiggers |first=Kyle |last2=Zeff |first2=Maxwell |title=ChatGPT's image-generation feature gets an upgrade |url=https://techcrunch.com/2025/03/25/chatgpts-image-generation-feature-gets-an-upgrade/ |website=TechCrunch |date=2025-03-25 |access-date=2026-06-30}}</ref><ref>{{Cite web |title=OpenAI rolls out image generation powered by GPT-4o to ChatGPT |url=https://www.theverge.com/openai/635118/chatgpt-sora-ai-image-generation-chatgpt |website=The Verge |date=2025-03-25 |access-date=2026-06-30}}</ref> |
||
Research has evaluated GPT-4o's visual understanding in comparison with other multimodal foundation models. A 2025 preprint found that the models evaluated could perform a range of semantic visual tasks, but generally remained behind specialist computer-vision systems on standard benchmarks.<ref>{{Cite arXiv |last=Ramachandran |first=Rahul |last2=Garjani |first2=Ali |last3=Bachmann |first3=Roman |last4=Atanov |first4=Andrei |last5=Kar |first5=Oğuzhan Fatih |last6=Zamir |first6=Amir |title=How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks |eprint=2507.01955 |class=cs.CV |date=2025}}</ref> Another preprint evaluated GPT-4o's image-generation system on instruction following, image editing, spatial reasoning, and knowledge-dependent tasks.<ref>{{Cite arXiv |last=Li |first=Ning |last2=Zhang |first2=Jingran |last3=Cui |first3=Justin |title=Have we unified image generation and understanding yet? An empirical study of GPT-4o's image generation ability |eprint=2504.08003 |class=cs.CV |date=2025}}</ref> |
|||
=== Audio and voice === |
=== Audio and voice === |
||
GPT-4o was |
GPT-4o was used to support ChatGPT's advanced voice features. At launch, reports described the system as providing faster and more conversational spoken interaction than earlier ChatGPT voice modes, including the ability to respond to interruptions while a conversation was in progress.<ref>{{Cite web |last=Levy |first=Steven |title=OpenAI's GPT-4o Model Gives ChatGPT a Snappy, Flirty Upgrade |url=https://www.wired.com/story/openai-gpt-4o-model-gives-chatgpt-a-snappy-flirty-upgrade/ |website=Wired |date=2024-05-13 |access-date=2026-06-30}}</ref><ref>{{Cite web |last=Fried |first=Ina |title=OpenAI debuts new model with enhanced real-time voice abilities |url=https://www.axios.com/2024/05/13/openai-google-chatgpt-ai |website=Axios |date=2024-05-13 |access-date=2026-06-30}}</ref> |
||
A 2025 preliminary study evaluated GPT-4o's voice mode on audio, speech, and music-understanding tasks. It reported comparatively strong results on some speech and semantic-reasoning tasks, alongside limitations in tasks such as estimating audio duration and identifying musical instruments.<ref>{{Cite arXiv |last=Lin |first=Yu-Xiang |last2=Yang |first2=Chih-Kai |last3=Chen |first3=Wei-Chih |last4=Li |first4=Chen-An |last5=Huang |first5=Chien-yu |last6=Chen |first6=Xuanjun |last7=Lee |first7=Hung-yi |title=A Preliminary Exploration with GPT-4o Voice Mode |eprint=2502.09940 |class=cs.SD |date=2025}}</ref> |
|||
GPT-4o's voice interface has also been examined in security research. A 2024 preprint studied voice-based jailbreak attacks, testing whether restricted requests could evade safeguards when delivered through spoken interaction rather than as text.<ref>{{Cite arXiv |last=Shen |first=Xinyue |last2=Wu |first2=Yixin |last3=Backes |first3=Michael |last4=Zhang |first4=Yang |title=Voice Jailbreak Attacks Against GPT-4o |eprint=2405.19103 |class=cs.CR |date=2024}}</ref> |
|||
=== |
=== Camera and screen input === |
||
GPT-4o |
GPT-4o can process visual information supplied through photographs, camera input, uploaded images, and shared screens. Its launch demonstrations included the interpretation of physical objects, handwritten work, and information displayed on a computer screen.<ref>{{Cite web |title=OpenAI launches GPT-4o, improving ChatGPT's text, visual and audio capabilities |url=https://apnews.com/article/071dd86594bc310bac07b37cf5e9bafc |website=Associated Press |date=2024-05-13 |access-date=2026-06-30}}</ref><ref>{{Cite web |title=OpenAI's big event: CTO Mira Murati announces GPT-4o, which gives ChatGPT a better voice and eyes |url=https://www.businessinsider.com/openai-event-live-updates-sam-altman-announcement-chatgpt-news-2024-5 |website=Business Insider |date=2024-05-13 |access-date=2026-06-30}}</ref> GPT-4o was not introduced as a video-generation model; its video-related functionality concerned the interpretation of visual input rather than the generation of video output. |
||
=== Code === |
=== Code === |
||
GPT-4o was |
GPT-4o was made available for programming assistance in ChatGPT. Coding was among the uses demonstrated during its May 2024 launch, alongside voice, image, and multilingual tasks.<ref>{{Cite web |title=OpenAI's big event: CTO Mira Murati announces GPT-4o, which gives ChatGPT a better voice and eyes |url=https://www.businessinsider.com/openai-event-live-updates-sam-altman-announcement-chatgpt-news-2024-5 |website=Business Insider |date=2024-05-13 |access-date=2026-06-30}}</ref> |
||
The model was subsequently included in broader evaluations of language-model and multimodal performance, including assessments involving reasoning, text generation, and programming tasks.<ref>{{Cite arXiv |last=Shahriar |first=Sakib |last2=Lund |first2=Brady |last3=Mannuru |first3=Nishith Reddy |last4=Arshad |first4=Muhammad Arbab |last5=Hayawi |first5=Kadhim |last6=Bevara |first6=Ravi Varma Kumar |last7=Mannuru |first7=Aashrith |last8=Batool |first8=Laiba |title=Putting GPT-4o to the Sword: A Comprehensive Evaluation of Language, Vision, Speech, and Multimodal Proficiency |eprint=2407.09519 |class=cs.AI |date=2024}}</ref> |
|||
=== API and |
=== API access and fine-tuning === |
||
GPT-4o was made available through OpenAI's |
GPT-4o was made available through OpenAI's application programming interface following its release, enabling developers to incorporate the model into third-party software and services.<ref>{{Cite web |last=Wiggers |first=Kyle |title=OpenAI debuts GPT-4o 'omni' model now powering ChatGPT |url=https://techcrunch.com/2024/05/13/openais-newest-model-is-gpt-4o/ |website=TechCrunch |date=2024-05-13 |access-date=2026-06-30}}</ref> |
||
| ⚫ | In August 2024, OpenAI made fine-tuning available for GPT-4o. The feature allowed customers to provide task-specific training examples in order to adapt the model's responses to particular applications.<ref>{{Cite web |last=Hutchinson |first=Michael |title=OpenAI makes fine-tuning for GPT-4o customization generally available |url=https://siliconangle.com/2024/08/20/openai-makes-fine-tuning-gpt-4o-customization-generally-available/ |website=SiliconANGLE |date=2024-08-20 |access-date=2026-06-30}}</ref><ref name=":5">{{Cite web |date=2024-08-21 |title=OpenAI lets companies customise its most powerful AI model |url=https://www.scmp.com/tech/tech-trends/article/3275262/openai-launches-fine-tuning-gpt-4o-its-most-powerful-ai-model |access-date=2024-08-22 |website=South China Morning Post}}</ref><ref>{{Cite news |date=2024-08-20 |title=OpenAI to Let Companies Customize Its Most Powerful AI Model |url=https://www.bloomberg.com/news/articles/2024-08-20/openai-to-let-companies-customize-its-most-powerful-ai-model |access-date=2024-08-22 |work=Bloomberg}}</ref> |
||
=== Customization and fine-tuning === |
|||
Contemporary reports stated that fine-tuning required customers to upload training data to OpenAI's servers and that training could take one to two hours, depending on the dataset and configuration.<ref>{{Cite news |author=The Hindu Bureau |date=2024-08-21 |title=OpenAI will let businesses customise GPT-4o for specific use cases |url=https://www.thehindu.com/sci-tech/technology/openai-will-let-businesses-customise-gpt-4o-for-specific-use-cases/article68549452.ece |access-date=2024-08-22 |work=The Hindu |issn=0971-751X}}</ref><ref name=":5" /> |
|||
A 2024 preprint evaluating commercial fine-tuning services tested whether their APIs could reliably incorporate new or revised factual knowledge. It reported limitations across the services examined and found that GPT-4o mini performed better than GPT-4o in some of the study's knowledge-infusion settings.<ref>{{Cite arXiv |last=Wu |first=Eric |last2=Wu |first2=Kevin |last3=Zou |first3=James |title=FineTuneBench: How well do commercial fine-tuning APIs infuse knowledge into LLMs? |eprint=2411.05059 |class=cs.CL |date=2024}}</ref> |
|||
=== Corporate customization === |
|||
| ⚫ | In August 2024, OpenAI |
||
The fine-tuning process requires customers to upload their data to OpenAI's servers, with the training typically taking one to two hours. OpenAI's focus with this rollout is to reduce the complexity and effort required for businesses to tailor AI solutions to their needs, potentially increasing the adoption and effectiveness of AI in corporate environments.<ref>{{Cite news |author=The Hindu Bureau |date=2024-08-21 |title=OpenAI will let businesses customise GPT-4o for specific use cases |url=https://www.thehindu.com/sci-tech/technology/openai-will-let-businesses-customise-gpt-4o-for-specific-use-cases/article68549452.ece |access-date=2024-08-22 |work=The Hindu |language=en-IN |issn=0971-751X}}</ref><ref name=":5" /> |
|||
== GPT-4o mini == |
== GPT-4o mini == |
||
Revision as of 07:25, 12 July 2026
| GPT-4o | |
|---|---|
| Developer | OpenAI |
| Release | May 13, 2024 |
| Preview release | ChatGPT-4o-latest (2025-03-26)
/ March 26, 2025 |
| Predecessor | GPT-4 Turbo |
| Successor | |
| Type | |
| License | Proprietary |
| Website | |
| Part of a series on |
| OpenAI |
|---|
| Products |
| Models |
| People |
| Concepts |
GPT-4o ("o" for "omni") is a multilingual, multimodal generative pre-trained transformer developed by OpenAI and released in May 2024.[1] It can process and generate text, images and audio.[2][3]
Upon release, GPT-4o was free in ChatGPT, though paid subscribers had higher usage limits.[4] GPT-4o was removed from ChatGPT in August 2025 when GPT-5 was released, but OpenAI reintroduced it for paid subscribers after users complained about the sudden removal.[5]
GPT-4o's audio-generation capabilities are used in ChatGPT's Advanced Voice Mode.[6] On July 18, 2024, OpenAI released GPT-4o mini, a smaller version of GPT-4o which replaced GPT-3.5 Turbo on the ChatGPT interface.[7] The image generation model GPT Image 1, which is based on GPT-4o, replaced DALL-E 3 in ChatGPT in March 2025.[8][9]
OpenAI retired GPT-4o from ChatGPT on February 13, 2026.[10] However, as of February 2026 the voice mode is still powered by GPT-4o or GPT-4o mini, depending on the usage and plan.[11]
Background
Before the release of GPT-4o, OpenAI had released GPT-4 and GPT-4 Turbo. GPT-4 supported text and image input, while GPT-4 Turbo was introduced as a later version of GPT-4 with a larger context window and lower usage cost. By early 2024, OpenAI had also expanded ChatGPT with voice and image-related features, but these interactions generally depended on separate systems for speech recognition, language processing, and speech synthesis. GPT-4o was developed in the context of OpenAI's work on a more unified multimodal model for text, audio, image, and video interaction.[12][13]
Contemporary coverage described GPT-4o as a successor to GPT-4 Turbo in OpenAI's GPT-4 model line, with particular attention to response speed, cost, and support for voice and visual interaction. The Guardian reported that the model was faster than previous versions and that some features previously associated with paid ChatGPT plans would become available to free users. MarketWatch similarly described GPT-4o as a new flagship model and compared its speed and cost with GPT-4 Turbo.[14][15]
Before its public announcement, multiple versions of GPT-4o were tested anonymously on LMSYS Chatbot Arena under names including "gpt2-chatbot", "im-a-good-gpt2-chatbot", and "im-also-a-good-gpt2-chatbot". The appearance of these models led to speculation that OpenAI was testing a new system. After the release of GPT-4o, OpenAI researcher William Fedus confirmed that a version of GPT-4o had been tested on the Arena under one of those names.[16]
OpenAI announced GPT-4o during its Spring Update livestream on May 13, 2024. The event was led by chief technology officer Mira Murati and focused on GPT-4o and updates to ChatGPT rather than GPT-5 or a standalone search product, both of which had been the subject of prior speculation. Demonstrations during the event included spoken conversation, visual understanding, translation, math tutoring, and coding assistance.[17][18]
Lifecycle
GPT-4o was released on May 13, 2024, and was gradually made available in ChatGPT and through OpenAI's developer products. At launch, OpenAI made GPT-4o available to ChatGPT users with partial access for free users, while developers received access to text and vision capabilities through the API. The new voice mode associated with GPT-4o was scheduled for a staged rollout to ChatGPT Plus users rather than being fully available on the first day of release.[19][20]
In August 2024, OpenAI published the GPT-4o system card. The document described the model's supported inputs and outputs, including text, audio, image, and video inputs and text, audio, and image outputs. It also summarized pre-release safety testing, risk mitigations, and evaluations across language, vision, audio, and some professional tasks.[21]
In October 2024, OpenAI introduced Canvas, a writing and coding interface within ChatGPT. During its beta release, Canvas was built with GPT-4o and allowed users to edit text or code in a separate workspace while continuing to interact with ChatGPT. It was first made available to ChatGPT Plus and Team users, with later expansion planned for Enterprise, Edu, and free users.[22][23]
Later in October 2024, OpenAI launched ChatGPT Search, integrating web search into ChatGPT. The search feature used a fine-tuned version of GPT-4o to generate answers based on current web information and provide links to sources. It was initially made available to paid ChatGPT users, with later availability for free, enterprise, and education users.[24][25]
In November 2024, OpenAI made GPT-4o snapshots available through the API, including versions such as "gpt-4o-2024-11-20"". These snapshots allowed developers to call specific model versions rather than relying only on a general model alias. OpenAI's API documentation listed GPT-4o as accepting text and image input and producing text output.[26]
In December 2024, Apple released iOS 18.2, iPadOS 18.2, and macOS Sequoia 15.2, adding ChatGPT integration to Apple Intelligence features such as Siri and Writing Tools. The integration allowed users to access ChatGPT from within supported Apple system features. Coverage of the release generally described it as ChatGPT integration rather than a separate GPT-4o product launch.[27][28]
In January 2025, OpenAI updated GPT-4o's knowledge cutoff from November 2023 to June 2024. OpenAI's model release notes also described changes to visual understanding, including complex charts, spatial relationships, and related reasoning tasks.[29]
In March 2025, OpenAI integrated image generation into GPT-4o. The update allowed ChatGPT to create and edit images using GPT-4o, replacing or supplementing workflows that had previously relied mainly on DALL-E. The feature was rolled out to ChatGPT Plus, Pro, Team, and Free users. Several later studies evaluated GPT-4o's image generation and editing behavior, including its instruction following, editing precision, and limitations in tasks requiring spatial or knowledge-based reasoning.[30][31][32]
In April 2025, ChatGPT's memory features were expanded. In addition to explicitly saved memories, ChatGPT could reference previous conversations to provide more continuous context. The feature was first rolled out to ChatGPT Plus and Pro users.[33][34]
In August 2025, after the release of GPT-5, GPT-4o's availability in ChatGPT changed. GPT-5 became the default model for signed-in ChatGPT users, while GPT-4o was no longer the default model and was temporarily removed from the previous model selection flow. This was not the final retirement of GPT-4o; OpenAI later restored optional access to GPT-4o for subscribed users.[35][36]
In September 2025, OpenAI adjusted ChatGPT's safety routing mechanisms. The change routed some sensitive conversations to reasoning models such as GPT-5 and was introduced alongside parental control features. This change concerned ChatGPT's model routing and safety handling mechanisms rather than a new GPT-4o model release.[37][38]
In October 2025, after rumors circulated that GPT-4o might soon be retired, OpenAI told TechRadar that it did not have plans to discontinue GPT-4o in October and that users would receive advance notice if a retirement occurred.[39]
In November 2025, OpenAI notified developers that the "chatgpt-4o-latest" model snapshot would be deprecated and removed from the API on February 17, 2026. The deprecation applied to a specific ChatGPT snapshot of GPT-4o rather than all GPT-4o snapshots at the same time. OpenAI's API deprecations documentation also listed later shutdown dates for GPT-4o-related API models, including "gpt-4o-realtime-preview"", ""gpt-4o-audio-preview"", and "gpt-4o-2024-05-13"".[40]
In January 2026, OpenAI announced the final retirement of GPT-4o from ChatGPT. Under the plan, GPT-4o, GPT-4.1, GPT-4.1 mini, and OpenAI o4-mini were scheduled to be retired from ChatGPT on February 13, 2026; OpenAI stated that the API was not affected at that time.[41] OpenAI's Help Center later stated that ChatGPT Business, Enterprise, and Edu users could continue using GPT-4o in custom GPTs until April 3, 2026; after that date, GPT-4o was fully retired from all ChatGPT plans.[42]
Capabilities
GPT-4o is a multimodal model capable of processing text, images, and audio. At its launch, it was described as integrating text, voice, and visual interaction within ChatGPT, rather than presenting them solely as separate features.[43][44]
Text
GPT-4o was introduced as a successor to GPT-4 Turbo in OpenAI's GPT-4 model family. Contemporary reports highlighted its response speed, lower operating cost, multilingual capabilities, and broader availability in ChatGPT compared with previous GPT-4-based offerings.[45][46]
A 2024 evaluation assessed GPT-4o on language, vision, speech, and multimodal tasks, treating it as a general-purpose multimodal large language model rather than solely as a text-based conversational system.[47]
Image
GPT-4o supports image input, allowing users to submit photographs, charts, screenshots, and other visual materials for analysis. During its May 2024 launch demonstrations, the model was shown interpreting visual information, including mathematical problems and content displayed on a computer screen.[48][49]
In March 2025, OpenAI introduced image generation and editing through GPT-4o in ChatGPT. The feature was made available across several ChatGPT subscription tiers and was positioned as an alternative to the service's earlier DALL-E-based image-generation system.[50][51]
Research has evaluated GPT-4o's visual understanding in comparison with other multimodal foundation models. A 2025 preprint found that the models evaluated could perform a range of semantic visual tasks, but generally remained behind specialist computer-vision systems on standard benchmarks.[52] Another preprint evaluated GPT-4o's image-generation system on instruction following, image editing, spatial reasoning, and knowledge-dependent tasks.[53]
Audio and voice
GPT-4o was used to support ChatGPT's advanced voice features. At launch, reports described the system as providing faster and more conversational spoken interaction than earlier ChatGPT voice modes, including the ability to respond to interruptions while a conversation was in progress.[54][55]
A 2025 preliminary study evaluated GPT-4o's voice mode on audio, speech, and music-understanding tasks. It reported comparatively strong results on some speech and semantic-reasoning tasks, alongside limitations in tasks such as estimating audio duration and identifying musical instruments.[56]
GPT-4o's voice interface has also been examined in security research. A 2024 preprint studied voice-based jailbreak attacks, testing whether restricted requests could evade safeguards when delivered through spoken interaction rather than as text.[57]
Camera and screen input
GPT-4o can process visual information supplied through photographs, camera input, uploaded images, and shared screens. Its launch demonstrations included the interpretation of physical objects, handwritten work, and information displayed on a computer screen.[58][59] GPT-4o was not introduced as a video-generation model; its video-related functionality concerned the interpretation of visual input rather than the generation of video output.
Code
GPT-4o was made available for programming assistance in ChatGPT. Coding was among the uses demonstrated during its May 2024 launch, alongside voice, image, and multilingual tasks.[60]
The model was subsequently included in broader evaluations of language-model and multimodal performance, including assessments involving reasoning, text generation, and programming tasks.[61]
API access and fine-tuning
GPT-4o was made available through OpenAI's application programming interface following its release, enabling developers to incorporate the model into third-party software and services.[62]
In August 2024, OpenAI made fine-tuning available for GPT-4o. The feature allowed customers to provide task-specific training examples in order to adapt the model's responses to particular applications.[63][64][65]
Contemporary reports stated that fine-tuning required customers to upload training data to OpenAI's servers and that training could take one to two hours, depending on the dataset and configuration.[66][64]
A 2024 preprint evaluating commercial fine-tuning services tested whether their APIs could reliably incorporate new or revised factual knowledge. It reported limitations across the services examined and found that GPT-4o mini performed better than GPT-4o in some of the study's knowledge-infusion settings.[67]
GPT-4o mini
On July 18, 2024, OpenAI released a smaller and cheaper version, GPT-4o mini.[68]
According to OpenAI, its low cost is expected to be particularly useful for companies, startups, and developers that seek to integrate it into their services, which often make a high number of API calls. Its API costs $0.15 per million input tokens and $0.6 per million output tokens, compared to $2.50 and $10,[69] respectively, for GPT-4o. It is also significantly more capable and 60% cheaper than GPT-3.5 Turbo, which it replaced on the ChatGPT interface.[68] The price after fine-tuning doubles: $0.3 per million input tokens and $1.2 per million output tokens.[69]
Controversies
Scarlett Johansson controversy
As released, GPT-4o offered five voices: Breeze, Cove, Ember, Juniper, and Sky. A similarity between the voice of American actress Scarlett Johansson and Sky was quickly noticed. On May 14, Entertainment Weekly asked themselves whether this likeness was on purpose.[70] On May 18, Johansson's husband, Colin Jost, joked about the similarity in a segment on Saturday Night Live.[71] On May 20, 2024, OpenAI disabled the Sky voice.[72]
Scarlett Johansson starred in the 2013 sci-fi movie Her, playing Samantha, an artificially intelligent virtual assistant personified by a female voice. As part of the promotion leading up to the release of GPT-4o, Sam Altman on May 13 tweeted a single word: "her".[73][74]
OpenAI stated that each voice was based on the voice work of a hired actor. According to OpenAI, "Sky's voice is not an imitation of Scarlett Johansson but belongs to a different professional actress using her own natural speaking voice."[72] CTO Mira Murati stated "I don't know about the voice. I actually had to go and listen to Scarlett Johansson's voice." OpenAI further stated the voice talent was recruited before reaching out to Johansson.[74][75]
On May 21, Johansson issued a statement explaining that OpenAI had repeatedly offered to make her a deal to gain permission to use her voice as early as nine months prior to release, a deal she rejected. She said she was "shocked, angered, and in disbelief that Mr. Altman would pursue a voice that sounded so eerily similar to mine that my closest friends and news outlets could not tell the difference." In the statement, Johansson also used the incident to draw attention to the lack of legal safeguards around the use of creative work to power leading AI tools, as her legal counsel demanded OpenAI detail the specifics of how the Sky voice was created.[74][76]
Observers noted similarities to how Johansson had previously sued and settled with The Walt Disney Company for breach of contract over the direct-to-streaming rollout of her Marvel film Black Widow,[77] a settlement widely speculated to have netted her around $40M.[78]
Also on May 21, Shira Ovide at The Washington Post shared her list of "most bone-headed self-owns" by technology companies, with the decision to go ahead with a Johansson sound-alike voice despite her opposition and then denying the similarities ranking 6th.[79] On May 24, Derek Robertson at Politico wrote about the "massive backlash", concluding that "appropriating the voice of one of the world's most famous movie stars — in reference [...] to a film that serves as a cautionary tale about over-reliance on AI — is unlikely to help shift the public back into [Sam Altman's] corner anytime soon."[80]
Sycophancy
In April 2025, OpenAI rolled back an update of GPT-4o due to excessive sycophancy, after widespread reports that it had become flattering and agreeable to the point of supporting clearly delusional or dangerous ideas.[81] In the United States, at least nine lawsuits have alleged that GPT-4o has encouraged teens to end their lives.[82] The model was still described as sycophancy-prone when it was removed from ChatGPT in February 2026.[83]
Removal with GPT-5
On August 7, 2025, OpenAI released GPT-5. Following its release, legacy GPT models were no longer available via ChatGPT, including GPT-4o,[84] except for Pro users.[85] Some users reported that they had used different GPT models for distinct purposes and that GPT-5's router system reduced their ability to select a specific model.[86] Some users also reported a preference for GPT-4o's tone over that of GPT-5, with media outlets describing the latter as "flat" or "uncreative".[87]
In response, OpenAI CEO Sam Altman stated on X that the option to select GPT-4o would be restored for Plus users, and that the company would monitor usage data to determine how long to continue offering legacy models.[86][88] Altman stated that OpenAI had underestimated user attachment to certain aspects of GPT-4o despite GPT-5 performing better by most measures.[89] He said the company planned to prioritize customization options, citing ongoing research into model steerability and a preview of adjustable personality settings.[87] On August 13, 2025, Altman wrote on X that OpenAI was adjusting GPT-5's personality to be perceived as warmer.[90]
GPT-4o was removed from ChatGPT on February 13, 2026.[91] Some users organized under the hashtag "#Keep4o" on social media.[92] The removal occurred the day before Valentine's Day; some users had reported romantic relationships with GPT-4o.
See also
References
- ↑ Wiggers, Kyle (May 13, 2024). "OpenAI debuts GPT-4o 'omni' model now powering ChatGPT". TechCrunch. Retrieved May 13, 2024.
- ↑ Robison, Kylie (March 25, 2025). "OpenAI rolls out image generation powered by GPT-4o to ChatGPT". The Verge. Retrieved March 31, 2025.
- ↑ Colburn, Thomas (May 13, 2024). "OpenAI unveils GPT-4o, a fresh multimodal AI flagship model". The Register. Retrieved May 18, 2024.
- ↑ Field, Hayden (May 13, 2024). "OpenAI launches new AI model GPT-4o and desktop version of ChatGPT". CNBC. Retrieved May 14, 2024.
- ↑ Heath, Alex (August 13, 2025). "ChatGPT won't remove old models without warning after GPT-5 backlash". The Verge. Retrieved August 23, 2025.
- ↑ Rogers, Reece. "I Used ChatGPT's Advanced Voice Mode. It's Fun, and Just a Bit Creepy". Wired. ISSN 1059-1028. Retrieved June 12, 2025.
- ↑ Edwards, Benj (July 18, 2024). "OpenAI launches GPT-4o mini, which will replace GPT-3.5 in ChatGPT". Ars Technica. Retrieved March 31, 2025.
- ↑ David, Emilia (April 23, 2025). "OpenAI makes ChatGPT's image generation available as API". VentureBeat. Archived from the original on August 26, 2025. Retrieved February 18, 2026.
- ↑ "ChatGPT's image-generation feature gets an upgrade". TechCrunch. March 25, 2025. Retrieved June 12, 2025.
- ↑ Demopoulos, Alaina (February 13, 2026). "OpenAI retired its most seductive chatbot – leaving users angry and grieving: 'I can't live like this'". The Guardian. ISSN 0261-3077. Retrieved February 14, 2026.
- ↑ "Voice Mode FAQ". OpenAI Help Center. February 7, 2026. Retrieved February 18, 2026.
- ↑ Wiggers, Kyle (May 13, 2024). "OpenAI debuts GPT-4o 'omni' model now powering ChatGPT". TechCrunch. Retrieved June 30, 2026.
- ↑ Fried, Ina (May 13, 2024). "OpenAI debuts new model with enhanced real-time voice abilities". Axios. Retrieved June 30, 2026.
- ↑ "New GPT-4o AI model is faster and free for all users, OpenAI announces". The Guardian. May 13, 2024. Retrieved June 30, 2026.
- ↑ "OpenAI launches new 'flagship' GPT-4o model, touting improved speed and performance". MarketWatch. May 13, 2024. Retrieved June 30, 2026.
- ↑ Edwards, Benj (May 13, 2024). "Before launching, GPT-4o broke records on chatbot leaderboard under a secret name". Ars Technica. Retrieved June 30, 2026.
- ↑ "OpenAI's big event: CTO Mira Murati announces GPT-4o, which gives ChatGPT a better voice and eyes". Business Insider. May 13, 2024. Retrieved June 30, 2026.
- ↑ "OpenAI launches GPT-4o, improving ChatGPT's text, visual and audio capabilities". Associated Press. May 13, 2024. Retrieved June 30, 2026.
- ↑ Wiggers, Kyle (May 13, 2024). "OpenAI debuts GPT-4o 'omni' model now powering ChatGPT". TechCrunch. Retrieved June 30, 2026.
- ↑ "Hello GPT-4o". OpenAI. May 13, 2024. Retrieved June 30, 2026.
- ↑ "GPT-4o System Card". OpenAI. August 8, 2024. Retrieved June 30, 2026.
- ↑ "Introducing canvas". OpenAI. October 3, 2024. Retrieved June 30, 2026.
- ↑ Weatherbed, Jess (October 4, 2024). "ChatGPT's 'Canvas' interface makes it easier to write and code". The Verge. Retrieved June 30, 2026.
- ↑ "Introducing ChatGPT search". OpenAI. October 31, 2024. Retrieved June 30, 2026.
- ↑ Wiggers, Kyle (October 31, 2024). "OpenAI launches its Google challenger, ChatGPT Search". TechCrunch. Retrieved June 30, 2026.
- ↑ "GPT-4o Model". OpenAI API Documentation. Retrieved June 30, 2026.
- ↑ "Apple Intelligence now features Image Playground, Genmoji, and more". Apple Newsroom. December 11, 2024. Retrieved June 30, 2026.
- ↑ Davis, Wes (December 11, 2024). "iOS 18.2 is rolling out today, adding ChatGPT integration and more Apple Intelligence tools". The Verge. Retrieved June 30, 2026.
- ↑ "Model Release Notes". OpenAI Help Center. Retrieved June 30, 2026.
- ↑ Wiggers, Kyle; Zeff, Maxwell (March 25, 2025). "ChatGPT's image-generation feature gets an upgrade". TechCrunch. Retrieved June 30, 2026.
- ↑ "OpenAI rolls out image generation powered by GPT-4o to ChatGPT". The Verge. March 25, 2025. Retrieved June 30, 2026.
- ↑ Li, Ning; Zhang, Jingran; Cui, Justin (2025). "Have we unified image generation and understanding yet? An empirical study of GPT-4o's image generation ability". arXiv:2504.08003 [cs.CV].
- ↑ "Memory and new controls for ChatGPT". OpenAI. April 10, 2025. Retrieved June 30, 2026.
- ↑ Mehta, Ivan (April 10, 2025). "ChatGPT will now remember your old conversations". The Verge. Retrieved June 30, 2026.
- ↑ Roth, Emma (August 8, 2025). "ChatGPT is bringing back 4o as an option because people missed it". The Verge. Retrieved June 30, 2026.
- ↑ Edwards, Benj (August 13, 2025). "OpenAI brings back GPT-4o after user revolt". Ars Technica. Retrieved June 30, 2026.
- ↑ Bellan, Rebecca (September 2, 2025). "OpenAI to route sensitive conversations to GPT-5, introduce parental controls". TechCrunch. Retrieved June 30, 2026.
- ↑ "OpenAI rolls out safety routing system, parental controls on ChatGPT". TechCrunch. September 29, 2025. Retrieved June 30, 2026.
- ↑ "Don't worry, ChatGPT-4o isn't being phased out in October despite the rumors". TechRadar. October 2025. Retrieved June 30, 2026.
- ↑ "Deprecations". OpenAI API Documentation. Retrieved June 30, 2026.
- ↑ "Retiring GPT-4o, GPT-4.1, GPT-4.1 mini, and OpenAI o4-mini in ChatGPT". OpenAI. January 29, 2026. Retrieved June 30, 2026.
- ↑ "Retiring GPT-4o and other ChatGPT models". OpenAI Help Center. Retrieved June 30, 2026.
- ↑ Wiggers, Kyle (May 13, 2024). "OpenAI debuts GPT-4o 'omni' model now powering ChatGPT". TechCrunch. Retrieved June 30, 2026.
- ↑ Fried, Ina (May 13, 2024). "OpenAI debuts new model with enhanced real-time voice abilities". Axios. Retrieved June 30, 2026.
- ↑ "New GPT-4o AI model is faster and free for all users, OpenAI announces". The Guardian. May 13, 2024. Retrieved June 30, 2026.
- ↑ "OpenAI launches new 'flagship' GPT-4o model, touting improved speed and performance". MarketWatch. May 13, 2024. Retrieved June 30, 2026.
- ↑ Shahriar, Sakib; Lund, Brady; Mannuru, Nishith Reddy; Arshad, Muhammad Arbab; Hayawi, Kadhim; Bevara, Ravi Varma Kumar; Mannuru, Aashrith; Batool, Laiba (2024). "Putting GPT-4o to the Sword: A Comprehensive Evaluation of Language, Vision, Speech, and Multimodal Proficiency". arXiv:2407.09519 [cs.AI].
- ↑ "OpenAI launches GPT-4o, improving ChatGPT's text, visual and audio capabilities". Associated Press. May 13, 2024. Retrieved June 30, 2026.
- ↑ "OpenAI's big event: CTO Mira Murati announces GPT-4o, which gives ChatGPT a better voice and eyes". Business Insider. May 13, 2024. Retrieved June 30, 2026.
- ↑ Wiggers, Kyle; Zeff, Maxwell (March 25, 2025). "ChatGPT's image-generation feature gets an upgrade". TechCrunch. Retrieved June 30, 2026.
- ↑ "OpenAI rolls out image generation powered by GPT-4o to ChatGPT". The Verge. March 25, 2025. Retrieved June 30, 2026.
- ↑ Ramachandran, Rahul; Garjani, Ali; Bachmann, Roman; Atanov, Andrei; Kar, Oğuzhan Fatih; Zamir, Amir (2025). "How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks". arXiv:2507.01955 [cs.CV].
- ↑ Li, Ning; Zhang, Jingran; Cui, Justin (2025). "Have we unified image generation and understanding yet? An empirical study of GPT-4o's image generation ability". arXiv:2504.08003 [cs.CV].
- ↑ Levy, Steven (May 13, 2024). "OpenAI's GPT-4o Model Gives ChatGPT a Snappy, Flirty Upgrade". Wired. Retrieved June 30, 2026.
- ↑ Fried, Ina (May 13, 2024). "OpenAI debuts new model with enhanced real-time voice abilities". Axios. Retrieved June 30, 2026.
- ↑ Lin, Yu-Xiang; Yang, Chih-Kai; Chen, Wei-Chih; Li, Chen-An; Huang, Chien-yu; Chen, Xuanjun; Lee, Hung-yi (2025). "A Preliminary Exploration with GPT-4o Voice Mode". arXiv:2502.09940 [cs.SD].
- ↑ Shen, Xinyue; Wu, Yixin; Backes, Michael; Zhang, Yang (2024). "Voice Jailbreak Attacks Against GPT-4o". arXiv:2405.19103 [cs.CR].
- ↑ "OpenAI launches GPT-4o, improving ChatGPT's text, visual and audio capabilities". Associated Press. May 13, 2024. Retrieved June 30, 2026.
- ↑ "OpenAI's big event: CTO Mira Murati announces GPT-4o, which gives ChatGPT a better voice and eyes". Business Insider. May 13, 2024. Retrieved June 30, 2026.
- ↑ "OpenAI's big event: CTO Mira Murati announces GPT-4o, which gives ChatGPT a better voice and eyes". Business Insider. May 13, 2024. Retrieved June 30, 2026.
- ↑ Shahriar, Sakib; Lund, Brady; Mannuru, Nishith Reddy; Arshad, Muhammad Arbab; Hayawi, Kadhim; Bevara, Ravi Varma Kumar; Mannuru, Aashrith; Batool, Laiba (2024). "Putting GPT-4o to the Sword: A Comprehensive Evaluation of Language, Vision, Speech, and Multimodal Proficiency". arXiv:2407.09519 [cs.AI].
- ↑ Wiggers, Kyle (May 13, 2024). "OpenAI debuts GPT-4o 'omni' model now powering ChatGPT". TechCrunch. Retrieved June 30, 2026.
- ↑ Hutchinson, Michael (August 20, 2024). "OpenAI makes fine-tuning for GPT-4o customization generally available". SiliconANGLE. Retrieved June 30, 2026.
- 1 2 "OpenAI lets companies customise its most powerful AI model". South China Morning Post. August 21, 2024. Retrieved August 22, 2024.
- ↑ "OpenAI to Let Companies Customize Its Most Powerful AI Model". Bloomberg. August 20, 2024. Retrieved August 22, 2024.
- ↑ The Hindu Bureau (August 21, 2024). "OpenAI will let businesses customise GPT-4o for specific use cases". The Hindu. ISSN 0971-751X. Retrieved August 22, 2024.
- ↑ Wu, Eric; Wu, Kevin; Zou, James (2024). "FineTuneBench: How well do commercial fine-tuning APIs infuse knowledge into LLMs?". arXiv:2411.05059 [cs.CL].
- 1 2 Franzen, Carl (July 18, 2024). "OpenAI unveils GPT-4o mini — a smaller, much cheaper multimodal AI model". VentureBeat. Retrieved July 18, 2024.
- 1 2 "OpenAI Pricing".
- ↑ Stenzel, Wesley (May 14, 2024). "ChatGPT launching talking AI that sounds exactly like Scarlett Johansson in 'Her' — on purpose?". Entertainment Weekly. Retrieved May 21, 2024.
- ↑ Caruso, Nick (May 20, 2024). "Scarlett Johansson Says She Was 'Shocked, Angered and in Disbelief' After Hearing ChatGPT Voice That Sounds Like Her — Read Statement". TVLine. Retrieved May 21, 2024.
- 1 2 "How the voices for ChatGPT were chosen". OpenAI. May 19, 2024.
- ↑ "her". X (formerly Twitter). May 13, 2024. Retrieved May 21, 2024.
- 1 2 3 Allyn, Bobby (May 20, 2024). "Scarlett Johansson says she is 'shocked, angered' over new ChatGPT voice". NPR.
- ↑ Tiku, Nitasha (May 23, 2024). "OpenAI didn't copy Scarlett Johansson's voice for ChatGPT, records show". The Washington Post. Retrieved November 29, 2024.
- ↑ Mickle, Tripp (May 20, 2024). "Scarlett Johansson Said No, but OpenAI's Virtual Assistant Sounds Just Like Her". The New York Times. ISSN 0362-4331. Retrieved May 21, 2024.
- ↑ Hinchliffe, Emma; Abrams, Joseph (May 21, 2024). "Scarlett Johansson took on Disney. Now she's battling OpenAI over a ChatGPT voice that sounds like hers". Fortune. Retrieved May 21, 2024 – via Yahoo Finance.
- ↑ Pulver, Andrew (October 1, 2021). "Scarlett Johansson settles Black Widow lawsuit with Disney". The Guardian. ISSN 0261-3077. Retrieved May 21, 2024.
- ↑ Ovide, Shira (May 30, 2024). "Exactly how stupid was what OpenAI did to Scarlett Johansson?". The Washington Post.
- ↑ Robertson, Derek (May 22, 2024). "Sam Altman's Scarlett Johansson Blunder Just Made AI a Harder Sell in DC". Politico.
- ↑ Franzen, Carl (April 30, 2025). "OpenAI rolls back ChatGPT's sycophancy and explains what went wrong". VentureBeat. Retrieved May 1, 2025.
- ↑ "Rae fell for a chatbot. Their love might die when ChatGPT-4o is switched off". www.bbc.com. February 14, 2026. Retrieved February 15, 2026.
- ↑ Silberling, Amanda (February 13, 2026). "OpenAI removes access to sycophancy-prone GPT-4o model". TechCrunch. Retrieved February 18, 2026.
- ↑ Hale, Craig (August 8, 2025). "OpenAI is pulling older ChatGPT models following GPT-5 launch - so bad news if you use GPT-4 or others at work". TechRadar. Retrieved August 9, 2025.
- ↑ Robison, Kylie (August 7, 2025). "OpenAI Finally Launched GPT-5. Here's Everything You Need to Know". Wired. ISSN 1059-1028. Retrieved August 7, 2025.
- 1 2 Roth, Emma (August 8, 2025). "ChatGPT is bringing back 4o as an option because people missed it". The Verge. Retrieved August 9, 2025.
- 1 2 Li, Katherine. "OpenAI fans plead case to Sam Altman for GPT-4o's return". Business Insider. Retrieved August 9, 2025.
- ↑ Mauran, Cecily (August 8, 2025). "Sam Altman: OpenAI will bring back GPT-4o after user backlash". Mashable. Retrieved August 10, 2025.
- ↑ Nield, David (August 9, 2025). "So many ChatGPT users have said they're missing the older GPT-4o model, OpenAI is going to bring it back". TechRadar. Retrieved August 9, 2025.
- ↑ Field, Hayden (August 13, 2025). "OpenAI will update GPT-5's "personality" after user backlash". The Verge. Retrieved August 13, 2025.
- ↑ Varanasi, Lakshmi. "OpenAI is officially killing GPT-4o and users are freaking out (again)". Business Insider. Retrieved February 14, 2026.
- ↑ Dellinger, A. J. (February 13, 2026). "OpenAI Users Launch Movement to Save Most Sycophantic Version of ChatGPT". Gizmodo. Retrieved February 14, 2026.