# Writer‘s Quest to Build an AI That Never Hallucinates

- Canonical: https://33rdsquare.com/writer-startup-launches-the-ai-model-which-never-hallucinates/
- Published: 2024-09-03
- Author: Jordan Brown
- Categories: [Artificial Intelligence & Machine Learning & ChatGPT](https://33rdsquare.com/category/tech/ai/)

---

In the burgeoning field of generative artificial intelligence, few challenges loom larger than the problem of AI "hallucinations"—those pesky instances where AI models spit out content that seems believable but is actually inconsistent with reality. Whether it‘s [GPT-3 concocting fake quotes from real people](https://www.technologyreview.com/2020/08/22/1007539/gpt3-openai-language-generator-artificial-intelligence-ai-opinion/) or [DALL-E Mini conjuring up photorealistic images of events that never happened](https://www.theverge.com/2022/6/6/23156948/ai-generated-art-dall-e-mini-images-ai-art), the capacity for AI to confabulate is both awe-inspiring and unnerving.

But what if there was a way to create AI models that could generate human-like text without the risk of detaching from the truth? That‘s the audacious goal Writer, a fast-rising AI startup, has set for itself with the release of its latest writing assistant powered by a model it calls "Palmyra." According to Writer‘s founders, this AI wordsmith is incapable of inserting falsehoods or unsupported claims into the content it generates. In other words, it‘s an AI that "never hallucinates."

## Solving the Achilles‘ Heel of Language Models

To appreciate what Writer is attempting to pull off, it helps to understand a bit about how today‘s state-of-the-art generative language models work under the hood. The core architecture behind systems like GPT-3 (which powers viral sensations like ChatGPT) consists of billions of parameters that have been trained on vast quantities of online data. Fed enough examples of human-written text, these models learn the statistical patterns and relationships between words, allowing them to "predict" the most probable next word in a sequence. String enough of these predictions together and you get remarkably coherent paragraphs that can be difficult to distinguish from text authored by a human.

The catch is that this probabilistic approach, while remarkably effective at producing fluent language, has no built-in mechanism for validating the factual accuracy of the model‘s outputs. As [one DeepMind researcher put it](https://www.technologyreview.com/2020/08/22/1007539/gpt3-openai-language-generator-artificial-intelligence-ai-opinion/), "Language models don‘t have a notion of truth; they just have a notion of what‘s plausible." In other words, GPT-3 and its ilk are not so much knowledge bases as "bullshit generators," to use a blunter term. They combine snippets of information in statistically compelling but often nonsensical ways, blurring the line between fact and fabrication.

For certain applications, like open-ended creative writing or worldbuilding, this sort of unconstrained imagination can be an asset. Who cares if the AI spins a yarn about [Barack Obama and his time at Hogwarts](https://www.technologyreview.com/2020/08/22/1007539/gpt3-openai-language-generator-artificial-intelligence-ai-opinion/)? But for enterprises looking to automate content workflows or provide AI writing assistance to employees, any deviation from established facts is a major liability. An AI that can wax poetic but can‘t be trusted to stick to the truth is not one you want writing your press releases or quarterly earnings reports.

This is the problem Writer aims to solve with Palmyra, the proprietary natural language generation (NLG) model at the heart of its platform. With 30 billion parameters, it‘s only slightly smaller than GPT-3 in terms of raw capacity. But Writer‘s founders claim they‘ve engineered a radically different kind of language model—one that can tap the expressive power of modern NLG without the "hallucination" problem that has hamstrung other systems.

## Baking Factuality Into the Model Architecture

So how does one go about building an AI writing assistant incapable of outputting false information? For Writer, the key is a novel architecture that combines several cutting-edge techniques in natural language processing (NLP):

1. **Encoder-Decoder Structure**: Unlike GPT-3 which uses a simple "decoder-only" structure to predict the next word based on previous context, Palmyra employs a more sophisticated "encoder-decoder" system. The encoder digests and compresses the input text into a rich numerical representation capturing its meaning, while the decoder takes that representation and generates an appropriate output. This allows for greater control and specificity in the outputs.
2. **Knowledge Grounding**: Before generating any text, Palmyra‘s encoder first consults an extensive knowledge base custom-built for each client. This includes documents like product specs, brand guidelines, financial reports, technical manuals, and other authoritative sources vetted for accuracy. By "grounding" the generation process in verified facts, the model is constrained to outputs aligned with established reality. Writer calls this a "path to explainability" since every assertion can be traced back to a specific trusted source.
3. **Fact Verification**: The model is also trained to cross-reference its planned response with the knowledge base before generating each word. If a claim cannot be substantiated by available facts, the decoder will not include it. This real-time fact-checking keeps the model honest and prevents it from "hallucinating" unsupported information.
4. **Uncertainty Quantification**: In cases where the model is asked a question it is not fully confident answering based on its ingested knowledge, it will express that uncertainty directly or refrain from responding rather than trying to guess or extrapolate. This is a way of signaling to users the boundaries of the system‘s knowledge and avoiding false pretenses of understanding.

Together, these architectural elements allow Writer to impose strict guardrails on its AI‘s generative capabilities without sacrificing fluency or flexibility. The result is a writing assistant that dynamically adapts to each user‘s context and communication style while steadfastly adhering to established facts. It‘s an approach Writer‘s founders liken to "baking factuality into the model" from the ground up rather than trying to "bolt it on" after the fact with post-processing filters or human oversight.

## A Different Vision for Enterprise AI

Writer‘s emphasis on factuality and traceability reflects a broader high-stakes vision for the role of generative AI in the enterprise. As language models grow more powerful and pervasive, so too does the impetus to deploy them in evermore consequential and sensitive domains like healthcare, finance, legal services, and government. The risks of an AI giving wrong medical advice, misstating a company‘s financials, or presenting false information in a courtroom are simply too high to abide any margin for error.

At the same time, enterprises are desperate to tap the productivity and creativity benefits of AI assistance without running afoul of regulatory compliance or brand safety guidelines. A 2022 survey by IBM found that [35% of companies are already using NLP solutions](https://filecache.mediaroom.com/mr5mr_ibmnews/190846/IBM%27s%20Global%20AI%20Adoption%20Index%202022.pdf), with another 42% planning to adopt the technology in the next year. But trust remains a major barrier, with 61% of IT professionals citing "accuracy of results" as a significant concern.

Writer believes its "never hallucinates" proposition addresses this trust gap head-on, opening up new frontiers for generative AI in business settings. Some potential use cases include:

- **Sales and Marketing**: Automatically generate product descriptions, web copy, email campaigns, and social media posts aligned with approved brand language and messaging.
- **Customer Support**: Provide 24/7 chatbot assistance to customers that reliably answers common questions and escalates edge cases to human agents.
- **Financial Analysis**: Produce quarterly earnings reports, investor updates, and market research that is exhaustively fact-checked against a company‘s financial data.
- **Legal Document Review**: Summarize and analyze contracts and legal documents, highlight key clauses and entities, and ensure consistency with established precedents and boilerplate language.
- **Medical Note-Taking**: Convert doctor-patient conversations into structured clinical notes that are automatically coded and compliant with medical billing and records standards.

In each case, Writer‘s AI would allow companies to automate rote communication tasks while ensuring accuracy and consistency—and without constant human oversight. Employees could focus on higher-order work while still having a personalized writing assistant at their fingertips whenever needed. Call it the "autocomplete for enterprise."

## The Outlook for Controllable AI

Writer is by no means the only player in the enterprise AI writing space. Incumbents like Microsoft and Google are also racing to integrate generative language capabilities into their productivity suites, while venture-backed startups like Jasper and Copy.ai have gained traction with their own iterations on AI content creation. Even OpenAI has begun exploring more constrained versions of GPT-3 to support business use cases.

But few if any have gone to the same lengths as Writer to bake factual accuracy and truthfulness into the very architecture of their models. In this sense, Writer is not just building a better mousetrap but reconceptualizing the problem altogether. Rather than trying to clean up the outputs of a general purpose language model after the fact, it aims to create a new class of "constitutionally constrained" AI systems that are incapable of generating false or unsupported information in the first place.

It‘s an approach that aligns with growing calls in the AI ethics community for [more "truthful and controllable" language models](https://venturebeat.com/ai/anthropic-introduces-constitutional-ai-to-help-ensure-models-behave-responsibly/). As these systems grow more powerful and are applied to higher stakes domains, the need for built-in safeguards and behavioral constraints becomes paramount. We can‘t always rely on humans to spot every incidental hallucination or fact-check every claim. The models themselves must be imbued with an unwavering commitment to the truth.

Of course, engineering this sort of "constitutional AI" is easier said than done. Language is inherently slippery and ambiguous, and even the most carefully curated knowledge base will have gaps and inconsistencies. There will always be edge cases where the model makes an unsupported inference or fails to grasp the nuance of a particular claim. Eliminating hallucinations entirely may be an unrealistic ideal.

But by making factuality a core design principle rather than an afterthought, Writer is charting a promising path forward for enterprise AI. One where we can harness the power of NLG to augment and accelerate human knowledge work without constantly worrying whether the outputs can be trusted. An AI writing assistant that is not just articulate but unassailable in its command of the facts.

In the end, the quest to build an AI that "never hallucinates" is not just about improving a single product or capability. It‘s about fundamentally rethinking what we want these systems to be and what role we want them to play in our organizations and society at large. Do we want AIs that are engaging but erratic storytellers, or do we want AIs that are reliable and responsible partners in pursuit of the truth? Writer is betting big on the latter—and challenging the rest of the industry to follow suit.

---

Source: [Writer‘s Quest to Build an AI That Never Hallucinates](https://33rdsquare.com/writer-startup-launches-the-ai-model-which-never-hallucinates/)
