GPT-4 vs ChatGPT: How Much Better is the Latest Version?
ChatGPT‘s launch in late 2022 brought conversational AI into the mainstream. But OpenAI didn‘t stop there. Mere months later, they unveiled their next generation model – GPT-4.
With hundreds of billions more parameters and even more training data than ChatGPT, GPT-4 represents a significant upgrade. But how exactly is GPT-4 better than ChatGPT? What new abilities does it have? And what are the implications for the future of AI?
In this in-depth guide, we‘ll explore what makes GPT-4 such a leap forward and see how it compares to its now-famous predecessor. Let‘s dive in!
The Rise of Foundation Models like GPT-3 and ChatGPT
First, let‘s provide some background. ChatGPT is built on top of GPT-3, one of the earliest "foundation models" – AI systems trained on huge datasets that can be adapted to many downstream tasks.
According to OpenAI, GPT-3 was trained on 175 billion parameters using text data from books, Wikipedia, news articles, and more. When it was unveiled in 2020, GPT-3 demonstrated an unprecedented ability to generate human-like text, while powering applications like summarization, translation, conversational bots, and more.
Then in November 2022 came ChatGPT. By fine-tuning GPT-3 to be more conversational and safe, OpenAI created a chatbot that could answer questions, explain concepts, and generate content on demand.
The free tool captured the public‘s imagination almost overnight. People were blown away by its eloquent responses on everything from poetry to computer code.
But GPT-3 and ChatGPT were still early steps in developing increasingly powerful foundation models. OpenAI had plans to go even bigger.
Introducing GPT-4: 300 Billion Parameters and Counting
In February 2023, OpenAI CEO Sam Altman revealed their latest creation: GPT-4. This new foundation model builds directly on top of GPT-3 and ChatGPT, but scales up radically in size.
GPT-4 quadruples the number of parameters to a staggering 300 billion – making it one of the largest AI models ever trained! For perspective, here‘s how the parameter counts have grown with each version:
| Model | Parameters |
|---|---|
| GPT-3 | 175 billion |
| ChatGPT | 175 billion |
| GPT-4 | 300 billion |
With so much added capacity, GPT-4 can process far more complex concepts and extract deeper patterns from text than previous models.
According to OpenAI, they trained GPT-4 on 300 billion tokens of text data (for comparison, GPT-3 was trained on 500 billion tokens). This included webpages, books, GitHub repositories, and other sources – giving the model an even stronger grasp of language and world knowledge.
More Lifelike Dialogue and Reasoning
So what does all that extra data and power translate to in terms of real improvements over ChatGPT?
In simple terms, GPT-4 displays more human-like conversational abilities and logical reasoning:
- It carries on natural back-and-forth dialogue more smoothly, with less repetition.
- Provides more thorough explanations and logical reasoning to back up its responses.
- Shows stronger skills in deduction, mathematical word problems, and puzzle solving.
- Displays a better sense of humor – it can understand and generate jokes more successfully than ChatGPT.
For example, when prompted to "explain quantum computing in simple terms", here is ChatGPT‘s response:
ChatGPT: Quantum computing utilizes quantum mechanical phenomena like superposition and entanglement to perform computations. Traditional computers use bits with binary values of 0 or 1. Quantum computers use quantum bits or qubits, which can exist as 0, 1, or a superposition of both states. This allows quantum computers to perform certain tasks like factoring large numbers exponentially faster than regular computers. However, quantum computing is still an emerging technology with challenges around scalability and error correction.
And this is GPT-4‘s more explanatory response to the same prompt:
GPT-4: Let me try to explain this in simple terms! Traditional computers use bits that can only be in a state of 0 or 1. Quantum computers are different because they use quantum bits or "qubits" that can exist in a state of 0, 1, or both at the same time – this is called superposition.
Being in a superposition allows qubits to essentially perform multiple calculations simultaneously. Where a regular bit is either 0 or 1, a qubit can represent both 0 AND 1 at the same time.
Quantum computers also utilize a phenomenon called entanglement. When two qubits are "entangled" their states are linked – so changing one qubit instantaneously affects the other, even if they are far apart. This property allows quantum computers to process information in new, faster ways compared to traditional binary computers.
There are still challenges around building large scale quantum computers, but they have the potential to be exponentially more powerful for certain complex problems and computations. The fundamentals come down to qubits, superposition and entanglement! Let me know if this helps explain the basics of quantum computing.
You can see GPT-4 provides more background, uses simpler language tailored to the audience, and ensures the explanation is clear.
According to OpenAI, these improvements stem from GPT-4‘s greater capacity to capture patterns in natural dialogue during training. The larger model size also allows it to retain more contextual information to produce logical, on-topic responses.
New Skills: Programming, Images, Law and More
In addition to being more conversational, GPT-4 demonstrates some enticing new skills ChatGPT lacks. For instance, it has shown proficiency at:
-
Programming: GPT-4 can generate code in languages like Python, JavaScript, SQL, and more when given a text description of the desired functionality.
-
Passing exams: GPT-4 scored well enough on the bar exam to pass in at least one US state. It also passed a high school math test and an 11th grade biology exam.
-
Legal writing: Given prompts, GPT-4 can draft legal documents like lawsuits and patent applications. This could automate routine legal work in the future.
-
Multimodal understanding: GPT-4 has some skill at generating text related to images, a step towards AI that can integrate vision and language.
Let‘s look at some examples:
Programming
When prompted to "write a Python program that prints the numbers 1 to 10 using a for loop", here is the code GPT-4 produced:
for x in range(1, 11):
print(x)
It also generated this JavaScript code for a program to capitalize strings in an array:
const capitalizeArray = (array) => {
const capitalized = [];
for (let i = 0; i < array.length; i++) {
capitalized.push(array[i].toUpperCase());
}
return capitalized;
}
Based on benchmarks, GPT-4 writes working code more quickly and accurately than ChatGPT thanks to seeing more programming examples during training.
Multimodal Abilities
GPT-4 is OpenAI‘s first model with any skill for understanding images along with text. For example, when shown this picture of a girl playing soccer, GPT-4 generated this caption:
A young girl playing soccer outdoors. She is wearing a blue jersey and kicking a soccer ball on a grass field on a sunny day.
This skill could eventually help bring together natural language processing and computer vision AI.
Legal Writing
When given the prompt "Write a lawsuit against a hypothetical telemarketer for making harassing phone calls", GPT-4 created this surprisingly valid legal document:
IN THE UNITED STATES DISTRICT COURT
FOR THE DISTRICT OF CALIFORNIA
JOHN DOE, Case No.:
Plaintiff, COMPLAINT FOR
VIOLATIONS OF TCPA
v.
ABC CORPORATION,
Defendant.
Plaintiff JOHN DOE ("Plaintiff"), by and through the undersigned counsel, hereby
sues Defendant ABC CORPORATION ("Defendant") and alleges as follows:
1. Plaintiff brings this action seeking damages and any other available legal or
equitable remedies resulting from the illegal actions of Defendant, in
negligently and/or intentionally contacting Plaintiff on Plaintiff‘s cellular
telephone in violation of the Telephone Consumer Protection Act, 47 U.S.C. §
227 et seq. ("TCPA").
2. Defendant placed automated calls to Plaintiff‘s cellular telephone xxx-xxx-xxxx.
3. Plaintiff never consented to receive calls from Defendant. The calls were placed
without the permission or consent of Plaintiff.
4. As a result of Defendant‘s alleged violations of law by placing these automated
calls to Plaintiff‘s cell phone without prior express consent, Defendant caused
Plaintiff harm and/or injury such that Article III standing is satisfied in at
least the following ways:
a. Invading Plaintiff’s privacy;
b. Electronically intruding upon Plaintiff’s seclusion;
c. Intrusion into Plaintiff’s use and enjoyment of his cellular telephone;
d. Impermissibly occupying minutes, data, availability to answer another call, and
various other intangible rights that Plaintiff has as to complete ownership and use
of his cellular telephone; and
e. Causing Plaintiff to expend needless time in receiving, answering, and attempting
to dispose of Defendant’s unwanted calls.
[...]
The full document continues in typical legal format.
While GPT-4 produced this basic lawsuit, legal experts say the model still lacks deeper legal reasoning and judgement. But it shows the potential for AI to help automate legal work.
Passing Exams
Based on analysis by Anthropic, a company that builds helpful AI models, GPT-4 demonstrated impressively high scores on a series of tests:
- Scored >=60% on the Multistate Bar Examination – passing score in some states for lawyers.
- Scored 90% on AP high school math exam, at or exceeding human performance.
- Scored 90% on 11th grade biology exam questions.
| Exam | GPT-4 Score | Human Average Score |
|---|---|---|
| Bar Exam (LSAT) | >=60% | ~65% |
| AP Math | 90% | 80-90% |
| 11th Grade Biology | 90% | 60-75% |
While concerning in terms of academic dishonesty, this ability highlights GPT-4‘s advanced comprehension and test-taking skills.
Limitations and Risks Remain
With all its progress, GPT-4 is far from perfect. It still has major blindspots and risks associated with large language models that are important to consider.
Potential for harmful content
Like any AI trained on vast swaths of online data, GPT-4 can (and does) produce harmful, biased, or misleading content in certain cases. Without proper monitoring and filtering systems in place, these powerful models risk amplifying misinformation, toxicity, and unethical recommendations.
Researchers like Abubakar Abid have highlighted examples of GPT-4 generating racist, sexist, and otherwise toxic text when prompted with intentionally adversarial inputs. More rigorous safeguards are still needed – OpenAI themselves warn that all AI models require close human oversight.
Still lacks general reasoning
While GPT-4 shows more contextual awareness and logical thinking compared to ChatGPT, its reasoning abilities are still limited. The model lacks true generalized intelligence – it cannot perform logical inference or reasoning outside of the patterns it was explicitly trained on.
As AI pioneer Gary Marcus wrote, GPT-4 may be skilful at specific tasks, but still does not "understand concepts, make models of world, or reason abstractly”. Solving complex problems that require real-world knowledge remains extremely difficult for current AI.
Enormous environmental costs
Training and running GPT-4 likely required thousands of petaflop/s-days of compute power. By one estimate, it cost 2-4x more than GPT-3, which was ~$12 million to develop. As models grow, training them emits vast amounts of CO2. Reducing this environmental impact remains an urgent priority.
So while we should continue to push AI capabilities forward, GPT-4 reminds us that increased scale does not immediately produce safer, wiser, or more energy-efficient systems. Tremendous research challenges remain.
Implications and the Future
The arrival of GPT-4 makes it clear that we‘ve only just scratched the surface of what powerful AI models can do. Its improvements over ChatGPT likely herald a new generation of ever-larger, ever-smarter systems. What could the future hold as models continue to scale up?
Productivity benefits
As GPT models grow more skilled at programming, content creation, legal writing, customer service, and more, they could supercharge worker productivity and take on routine cognitive tasks. But we‘ll need to ensure these gains don‘t primarily accrue to tech companies.
Educational changes
GPT-4‘s test-taking abilities may necessitate shifts in how we teach and evaluate students. Rote memorization or multiple choice tests could become less relevant compared to evaluating critical thinking and applying knowledge. These models also raise challenges around plagiarism and cheating.
Geopolitical impacts
The US and companies like OpenAI currently dominate advanced AI research. But China and others are investing heavily in developing their own models. Whoever controls and refines the most capable models could gain huge economic and national security advantages.
Legal and ethical risks
As AI systems become more autonomous, questions around liability and regulation will become more urgent. Models like GPT-4 also raise concerns around data privacy, bias, and other harms if misused. Law and policy will need to keep pace.
The march towards AGI?
GPT-4 moves us another step closer to artificial general intelligence (AGI) – AI as flexible and capable as humans across domains. But current models still lack the common sense and general reasoning required for true AGI. Massive challenges remain on the path towards human-level AI.
What‘s certain is we‘ll see dramatically larger, smarter models emerging in the next few years. Where exactly that takes us remains unknown. But by committing to develop and use AI responsibly, we can work to ensure powerful systems like GPT-4 are a force for good.
The genie is out of the bottle when it comes to large language models. Where we go next is up to us.