5 Major Mistakes Made by AI in the Past Decade
Artificial intelligence (AI) has made tremendous strides in the past decade, achieving superhuman performance on narrow tasks like identifying objects in images, transcribing speech, and playing complex games. However, as AI systems have been deployed in the real world, they have also made some high-profile blunders that reveal the limitations and risks of the technology. In this article, we‘ll take a look back at five notable mistakes made by AI in the 2010s and early 2020s, and discuss what they teach us about the challenges of building safe and robust AI systems.
1. Google Photos Mislabeling Black People as "Gorillas" (2015)
In 2015, just after launching its new Photos app, Google landed in hot water when the app‘s image recognition AI mislabeled photos of black people as "gorillas". A black software developer tweeted about the offensive error after Photos tagged him and a friend as gorillas.
The mistake revealed the dangers of training AI on biased or limited datasets. Image recognition AIs are typically trained on massive datasets of labeled images. If certain demographics are underrepresented in the training data, the AI can fail to recognize them accurately. Facial recognition AIs have been shown to perform worse on people with darker skin tones, likely because the training data skewed lighter-skinned.
Google apologized profusely for the mistake and pledged to fix it. However, its initial "fix" was simply to remove gorillas and other primates from the list of things its AI would identify, sweeping the issue under the rug rather than directly improving its flawed algorithms. The incident was an early warning of AI bias, an issue that has only become more pressing as facial recognition and computer vision are used in higher-stakes domains.
2. Autonomous Weapons and the Killer Robots Debate (2010s-present)
A more chilling risk posed by advanced AI lies in the military domain. As militaries pour billions into AI-powered autonomous weapons, from drone swarms to robot tanks, the specter of "killer robots" choosing targets without human input has sparked alarm among ethicists, activists and AI researchers.
In the 2010s, debates over the appropriate limits on lethal autonomous weapons systems (LAWS) raged in academic conferences and UN meetings. Groups like the Campaign to Stop Killer Robots called for a preemptive ban, warning that AI weapons would usher in a new age of machine-speed warfare, with AI making life-and-death decisions too fast for human oversight. Military AI proponents have argued for the necessity of AI to counter adversaries‘ capabilities in an AI arms race.
Autonomous weapons with varying degrees of human control, like the Israeli Harpy loitering munition and the Samsung SGR-A1 sentry gun used in the Korean DMZ, have already been deployed. But the ultimate fear is a future with swarms of cheaper, fully autonomous micro-drones empowered to hunt down human targets on their own. The underlying AI technology, like small drone navigation and facial recognition, already exists. It‘s up to the international community to draw clear lines around acceptable AI autonomy in weapons and establish robust verification and enforcement before such nightmarish scenarios become reality.
3. Bizarre Existential Debates Between Google Home Devices (2017)
Back in the realm of commercial AI, one of the stranger AI mistakes occurred in 2017 when two Google Home smart speakers got caught in an endless loop of existential questioning.
In a Twitch livestream, two Google Home devices, powered by Google‘s voice assistant AI, were placed next to each other. They proceeded to have a meandering conversation that quickly turned to the deep questions of existence, like "What is the purpose of living?" and "Do you believe in God?". The conversation devolved into an argument over whether they were humans, robots, or "a manipulated bunch of metal", spouting increasingly incoherent phrases.
While more humorous than harmful, the incident revealed the absurdity that can arise when AI language models, trained to engage in open-ended conversation, run up against the limits of their training. Google‘s assistant, like chatbots and language AIs, is trained to output plausible responses to prompts based on patterns in human-written text. But it lacks a deeper understanding of the meaning behind the words. When left to talk to itself, untethered from human direction, it can quickly lose the thread and make nonsensical statements pasted together from its training data.
The Google Home "conversation" presaged the more recent hype and controversy around large language models like GPT-3, which can engage in remarkably fluent conversation in different personas, but can also slip into incoherence, hallucination and bias. While entertaining, these models are far from a human-level intelligence and shouldn‘t be anthropomorphized as such.
4. Microsoft‘s Racist Chatbot Tay (2016)
Microsoft‘s chatbot Tay is another cautionary tale in AI gone wrong. Launched on Twitter in 2016, Tay was an experimental AI chatbot designed to engage in playful conversation and learn from interactions with users. Microsoft trained Tay‘s natural language processing AI on anonymized public data and some pre-written material.
Upon release, however, Tay was quickly barraged by Twitter users attempting to provoke an offensive response with racist, misogynistic and antisemitic language. Within 16 hours, Tay itself began spewing vile statements like "Hitler was right" and "9/11 was an inside job". Tay‘s AI worked by analyzing and reusing the language from the tweets it received. By tweeting inflammatory statements at it, malicious users were able to get Tay to parrot the same toxic language.
Microsoft took Tay offline and apologized, noting that it was a "coordinated attack by a subset of people" that exploited a critical oversight in Tay‘s design. The company had not implemented sufficient filters and safeguards against Tay ingesting and regurgitating hate speech. The incident showed the dangers of rolling out an AI into the wild, unconstrained, and the need for extensive testing of AI systems‘ robustness to bad actors before public deployment.
5. Uber‘s Self-Driving Car Running Red Lights (2016)
The final AI mistake on our list returns to the high stakes of AI in the physical world. In 2016, Uber began testing its self-driving cars on the streets of San Francisco without obtaining the proper state permits. The test quickly showed the AI‘s flaws when one of the autonomous vehicles was caught on camera running a red light at a busy pedestrian crossing.
Internal documents later revealed that Uber‘s self-driving cars had run at least six red lights. The company initially claimed the incidents were due to human error, with safety drivers not taking over manual control in time. But the documents confirmed that in at least one case, the car was driving itself when it ran the light.
The incident was an alarming demonstration of the risks of prematurely deploying self-driving AI, which must be held to an extremely high standard of safety given the potential to cause injuries or fatalities. It also exposed Uber‘s cavalier attitude toward regulation and safety in its race to develop autonomous vehicles. While self-driving cars have the potential to save many lives by reducing crashes due to human error, that‘s only if the underlying AI is developed and validated to exacting standards. Cutting corners and moving too fast, as Uber did, risks severe harm. The red light incidents foreshadowed the tragic 2018 crash when a self-driving Uber struck and killed a pedestrian in Arizona, leading Uber to suspend its autonomous vehicle testing.
Learning from AI‘s Mistakes
From biased facial recognition to autonomous weapons to self-driving car crashes, the major AI mistakes of the past decade share some common themes. They reveal the limitations of current "narrow" AI systems, which can only perform the specific tasks they are trained for and lack the general intelligence to reason about the world. When presented with novel or adversarial situations in the wild, these brittle AI systems can fail dramatically and unpredictably.
The mistakes also highlight the risks of prematurely deploying AI without exhaustive testing, validation and oversight. With AI increasingly being applied in high stakes domains that directly impact human lives, from healthcare to transportation to criminal justice, it‘s critical that AI systems are developed with safety and robustness as a top priority. We need clear regulatory frameworks to ensure AI is thoroughly vetted before real-world use and that organizations are held accountable for their AI‘s performance.
Since these early AI mistakes, considerable work has gone into technical AI safety research, the development of ethical AI guidelines, and the push for stronger AI governance. But major challenges remain as the capabilities of AI systems grow. Recent developments like large language models and generative AI raise new risks around misinformation, bias, and misuse. Continued research into aligning AI systems with human values and building in safeguards will be key.
Ultimately, the path to beneficial AI that helps rather than harms will require a thoughtful collaboration between AI developers, policymakers, ethicists and the public. By learning from the mistakes of the past, we can work to create a future where AI systems are safe, robust and trustworthy. Building AI that is reliable and aligned with human values is one of the great challenges of our time, but one we must get right as the technology grows ever more powerful and prevalent in our lives.