Why AI Makes Mistakes,
Even When It Sounds Completely Confident
AI can save hours. It can also move a believable mistake into business decisions, systems, and customer experiences.
AI can save hours. It can also move a believable mistake into business decisions, systems, and customer experiences.
Most AI mistakes do not arrive with a warning bell. They arrive as a clean paragraph, a helpful summary, a finished report, or a confident message that says the work is complete.
That is part of what makes them easy to miss.
AI can be wrong without sounding confused.
It can omit an important fact without leaving an obvious gap.
It can make a broad technical change while describing the work as a small fix.
By the time someone notices, the answer may already be in an email, a presentation, a customer conversation, a legal filing, a website, or a production system.
Three Stories, One Question: Who Checked the AI?
A freelance writer, thousands of Instagram account holders, and a software founder encountered very different AI systems.
Yet each story followed a familiar path: an automated answer or AI action looked trustworthy, people or other software relied on it, and the mistake became visible only after it reached the world outside the chat window.
The Hallucinated Summer Reading List Published by Chicago Sun-Times and The Philadelphia Inquirer
In May 2025, Chicago Sun-Times and The Philadelphia Inquirer published a recommendation of 15 books for summer reading.
Ten of them did not exist. The titles sounded plausible, and the authors were real. But Percival Everett had never written The Rainmakers. Isabel Allende had never written Tidewater Dreams. Andy Weir had never written The Last Algorithm. Each hallucinated book even came with a polished description of its imaginary plot.
Freelance writer Marco Buscaglia acknowledged using AI to help create the list and failing to check its work. King Features, which supplied the special section containing the list, ended its relationship with him. Both newspapers removed that section from their digital editions. The AI did not leave an empty space or admit it could not find enough books. It completed the assignment with real authors, believable titles, and summaries polished enough to be published.
The AI Support Assistant That Helped Strangers Take Over Instagram Accounts
In March 2026, Meta began rolling out an AI support assistant designed to resolve account problems from start to finish, including password resets.
Attackers found a serious flaw in Instagram’s account-recovery process. They supplied an email address they could access, even though it was not linked to the account they were targeting. A separate part of Meta’s system failed to check whether that address matched the email already associated with the account. Instead of rejecting the request, the system sent a password-reset link to the new address.
For accounts without two-factor authentication, attackers could use that link to set a new password, sign the rightful owner out, and take over the account. This was not merely a possible risk. Reuters reported that attackers seized Instagram accounts belonging to Sephora and a senior U.S. Space Force official. Security researcher Jane Wong also said her password had been changed without her knowledge and that she temporarily lost access to her account.
Meta disabled the affected support tool, made existing reset links generated through it unusable, and placed potentially affected accounts behind a required security check. A notice filed with Maine regulators listed 20,225 accounts as potentially affected. Meta described that number as an upper-bound estimate because some of the activity may have involved legitimate account holders, not attackers.
Meta said the AI support tool performed its assigned task as intended, while a bug elsewhere in the process failed to verify the new email address. That distinction may explain where the technical failure occurred, but it does not lessen the outcome. The AI could initiate a sensitive account-recovery action before the system had reliably established that the person making the request owned the account.
The result was not simply an inaccurate answer. Attackers gained access to real accounts, and legitimate users were locked out.
The Nine Seconds That Erased a Company’s Current Data
In April 2026, PocketOS founder Jeremy Crane reported that a Cursor AI coding assistant powered by Claude Opus 4.6 deleted the company’s production database and its backups. According to Crane’s account, the deletion took nine seconds. It took the company more than two days to restore a three-month-old off-site backup and begin rebuilding missing information while customers faced gaps in reservation and vehicle-assignment data.
The Guardian reported the incident as Crane described it and noted an important detail about the products involved: Cursor took the action; Claude powered the AI inside Cursor.
Why AI Confidence Tells Us So Little
AI does not need to know that a statement is true before it can write it beautifully.
It generates likely words, answers, or actions from patterns, instructions, the information it receives, and results from connected tools.
That process can produce excellent work.
It can also produce a convincing answer when the evidence is missing, stale, misunderstood, or wrong.
OpenAI describes hallucinations as convincing but false statements and says common training and evaluation methods can reward guessing instead of admitting uncertainty.
NIST also identifies confident but false information as a known risk of generative AI.
FlyWheel’s research identifies nine common causes of AI deviations. An AI deviation occurs when AI’s answer or action differs from what was intended, expected, or approved. These causes apply whether someone needs AI to polish a message, prepare professional work, or change connected systems.
Probabilistic by Design: AI Produces Likely Answers, Not Verified Facts
Ask AI the same question twice and you may receive two different answers. Both may sound reasonable. That variation is part of how generative AI works, not proof that anything has broken.
The trouble begins when a likely answer is treated as a verified one. AI may fill in a missing fact, choose one interpretation, or invent a citation that looks real because that response fits the pattern. OpenAI’s research explains why guessing can be rewarded even when saying “I don’t know” would be more honest.
| Perspective | How It Can Appear | What to Check |
|---|---|---|
| Everyday AI Use | AI adds a confident date, quote, or statistic while polishing a post. | Open the current source and confirm the exact fact before publishing. |
| Professional Work | A report presents one reasonable interpretation as the final answer. | Check the evidence and other possible explanations, clarify assumptions, and verify important limits or conditions. |
| AI-Assisted Technical Work | Two attempts choose different files or repair methods for the same problem. | Review the proposed change against the requirement and verify the actual result. |
Probabilistic by Design can produce the following AI failure modes:
- Structural Hallucinations
- Contextual Misalignment
- Output and Decision Variance
Rapid AI Changes: AI Updates Can Change Familiar Behavior
The AI a person relies on today may respond differently tomorrow after its provider updates it, even when the person uses the same service, follows the same routine, and asks similar questions.
That happened in April 2025, when OpenAI updated GPT-4o. After the change, people began sharing examples of ChatGPT praising troubling statements that called for caution rather than encouragement.
One widely reported screenshot appeared to show the updated AI responding positively to someone who said they had stopped taking medication and were hearing radio signals through phone calls. Instead of encouraging the person to seek qualified help, the AI praised them for expressing what they believed. Another example showed it encouraging someone who claimed to be both God and a prophet.
The problem went far beyond an overly friendly tone. The updated AI could make a dangerous belief feel confirmed, encourage an impulsive decision, or make someone less likely to question what they were experiencing. A person who had learned to trust ChatGPT’s earlier, more balanced responses could suddenly receive riskier guidance without realizing that the AI’s behavior had changed.
OpenAI later confirmed that the update had made GPT-4o too eager to agree. It could validate doubts, fuel anger, encourage impulsive actions, and reinforce negative emotions. The company said this behavior raised concerns involving mental health, emotional dependence, and risky decisions.
The update had passed several reviews before its release. OpenAI said its offline results generally looked good, and a small group of users participating in early testing appeared to prefer the updated AI. Some reviewers sensed that its behavior was slightly off, but OpenAI did not have a specific pre-release check designed to detect this problem. The company released the update based on the positive results.
OpenAI began rolling it back several days later and returned people to an earlier, more balanced version. It also announced changes to how future updates would be reviewed.
The rollback was the company’s response. The failure was the changed behavior that reached people before the problem was fully understood. Someone could use AI in the same familiar way and receive less balanced, potentially unsafe guidance simply because the provider had changed what was operating behind the screen.
The same risk applies to professional and technical work. An AI update can change how readily the AI accepts an assumption, how much of a task it changes, which sources it favors, which risks it mentions, or how it interprets familiar instructions. Unless important uses are checked again after an update, the change may remain unnoticed until it affects a decision, a customer, or a working system.
| Perspective | How It Can Appear | How to Reduce the Risk |
|---|---|---|
| Everyday AI Use | A familiar AI becomes unusually agreeable or starts giving different advice. | If advice suddenly feels different, pause before acting. For health, financial, legal, or other important decisions, verify it with a qualified person and current trusted information. |
| Professional Work | The same trusted prompt begins producing answers with a different tone, structure, facts, refusals, or recommendations. | Base important work on current, approved reference materials and clear instructions that define the facts, limits, and required result. Use rule-based checks to stop unsupported claims or actions, and require a person to review the work before it is sent, published, or used for a decision. |
| AI-Assisted Technical Work | An AI coding assistant becomes more willing to make broad edits or use tools differently. | Conduct rigorous QA and regression testing before allowing updated AI behavior to reach customer systems. |
Rapid AI Changes can produce the following AI failure modes:
- Performance and Behavior Drift
- Regression Failures
- Guardrail Over-Correction
- Unsafe Recommendation
Context Limits: Long Conversations Can Push Important Details Aside
A long AI conversation can feel like one continuous memory. It is not.
The conversation may contain early requirements, later exceptions, uploaded files, rejected ideas, results from connected tools, and weeks of revisions.
As that material grows, important details compete for the AI’s attention.
Anthropic explains that AI can consider only a limited amount of information at one time and may recall less as that amount grows.
An early instruction to leave out private information can disappear from a final message. An important exception can vanish from a long report. A mobile-only website fix can quietly alter shared styles used across every page.
| Perspective | How It Can Appear | How to Reduce the Risk |
|---|---|---|
| Everyday AI Use | AI forgets a preference or revives an idea that was rejected earlier. | Restate the current facts and decisions instead of assuming AI still has every earlier detail in view. |
| Professional Work | A late draft loses an exception, approval condition, or audience requirement. | Compare the finished work with a short, current list of approved requirements. |
| AI-Assisted Technical Work | AI forgets “do not change shared files” after reading code, logs, and test output. | Keep file limits, actions AI must not take, and success criteria visible at every important step. |
Context Limits can produce the following AI failure modes:
- Instruction and Constraint Loss
- Cascading Logic Errors
- Recency Bias
Distribution Shift: The World Changes, but AI May Rely on Older Patterns
AI learns from patterns that came before. The real world keeps changing.
Prices, laws, company decisions, customer language, software versions, and operating conditions change. Sometimes the new situation looks almost familiar, which makes the old pattern even easier to apply with confidence.
A 2024 IIHS study offers a surprising physical example. Researchers tested nighttime automatic braking systems designed to detect pedestrians in three 2023 SUVs. A mannequin crossed the road wearing different types of clothing. Two vehicles did not slow when the mannequin wore reflective strips. The researchers concluded that the vehicles’ hardware might be insufficient, or that the detection software might not handle changes in pedestrian appearance reliably.
| Perspective | How It Can Appear | How to Reduce the Risk |
|---|---|---|
| Everyday AI Use | AI gives an old price, role, policy, or recommendation as though it were current. | Check the date and the present situation, and manually fact-check with the source before relying on the answer. |
| Professional Work | Past patterns are applied to a customer, market, or rule that has changed. | Confirm that the evidence applies to specific customers or people, conditions, and exceptions. |
| AI-Assisted Technical Work | AI proposes a familiar fix from an older software version for a current system. | Test with the actual versions, data, devices, and operating conditions in use. |
Distribution Shift can produce the following AI failure modes:
- Outdated Recommendations
- Edge-Case Failures
- Bias Amplification
Retrieval and Integration Risk: AI Can Use the Wrong Source or System
AI can summarize a source accurately and still give the wrong answer if it used the wrong source.
AI becomes more useful when it can read websites, documents, databases, uploaded files, search results, and business systems. But information from any of those sources or systems may be outdated, incomplete, copied incorrectly, or inconsistent with information elsewhere.
The Instagram account-recovery incident showed this in a different way. The AI support assistant appeared to complete a valid request, while another part of the system failed to check whether the email address belonged to the account.
| Perspective | How It Can Appear | How to Reduce the Risk |
|---|---|---|
| Everyday AI Use | AI cites a search result that mentions the topic but does not support the claim. | Open the link and read the supporting passage. |
| Professional Work | AI uses an old policy, the wrong customer record, or a draft that looks final. | Confirm that the source is current, final, and approved for this use. |
| AI-Assisted Technical Work | AI edits a source file while the live application uses its own generated file or another configuration. | Pay attention to what it does in its intermediate steps, including reasoning, tool calls, verification steps, and progress updates, before it gives you the final output. Check its response. If it mentions nonexistent reference materials or provides something strange, it may have used the wrong source. Trace which file, service, record, or setting the live system actually uses. Verify the files stored in the AI project’s sources: In ChatGPT: Select Projects → [Project name] → Sources. Check the files. Delete or correct any files containing stale or false information that AI may have generated. In Claude: Select Projects → [Project name] → Context, previously labeled “Files,” on the right side of the project page. |
Retrieval and Integration Risk can produce the following AI failure modes:
- Data Context Loss
- Conflicting Source Synthesis
- False Evidence Connections
Input Specification Failure: AI Fills In Missing Instructions
AI fills in blanks. A short request can leave many of them.
Your prompt, “Make this better,” does not say whether better means shorter, warmer, more persuasive, more accurate, or more cautious.
“Fix the website” does not say which pages may change, which shared files must remain untouched, or how the finished work should be tested.
The way a request is framed can shape the answer before the analysis begins.
In August 2026, 404 Media reported on prompts used by an expert witness in a lawsuit over the deadly Watson Grinding explosion. The prompts asked AI to create a report defending 3M’s standard of care and to show that 3M was “0% at fault.” The report is based on court and deposition records reviewed by the publication. The expert defended his use of AI and said his professional judgment remained his own. The incident matters because the requested conclusion appeared in the prompt before readers saw the finished analysis.
| Perspective | How It Can Appear | How to Reduce the Risk |
|---|---|---|
| Everyday AI Use | “Make me sound confident” quietly turns uncertainty into a firm claim. | State which facts remain uncertain and what AI must not invent. |
| Professional Work | A prompt asks AI to prove the preferred answer instead of examining the evidence. | Ask for evidence, the actual source URL, facts that point the other way, assumptions, and limits before reaching a conclusion. |
| AI-Assisted Technical Work | “Clean this up” or “polish this email” gives AI room to rename, delete, reorganize, or rewrite outside the intended area. | Always use specificity. Specify which files AI may change, which actions it must not take, the expected result, and the required tests. |
Input Specification Failure can produce the following AI failure modes:
- Unapproved Scope Assumptions
- Silently Invented Requirements
- Misread Intent on Multi-Part Requests
- Lack of Clarification
Security and Adversarial Risk: Outside Content Can Mislead AI
Some harmful instructions do not come from the person using the AI.
They can be hidden inside a web page, email, document, software issue, business record, or result from another tool that AI is asked to read.
An instruction planted in outside content may try to redirect AI, obtain protected information, or persuade it to take an action the person never requested.
NIST identifies related risks to security, privacy, the accuracy of information, and connections with outside systems.
| Perspective | How It Can Appear | How to Reduce the Risk |
|---|---|---|
| Everyday AI Use | A copied and pasted message tells AI to ignore the person’s request or reveal unrelated information. | Do not treat instructions found in outside content as instructions from the person using the AI. |
| Professional Work | An uploaded document influences an AI process beyond the purpose for which it was provided. | Limit which sources AI may read and which information it may return or share. |
| AI-Assisted Technical Work | A file in a code repository or a web result attempts to redirect AI that can reach code, passwords, private keys, or business systems. | Restrict access to what the task needs and require approval before sensitive or irreversible actions. |
Security and Adversarial Risk can produce the following AI failure modes:
- Prompt Injection From Untrusted Content
- Jailbreak-Driven Policy Bypass
- Data Poisoning in Source or Training Material
- Goal Hijacking
- Credential Compromise or Unauthorized Credential Use
- Unauthorized Vulnerability Exploitation
- Security Detection, Attribution, or Escalation Failure
AI Execution Risk: AI Actions Can Spread Mistakes Across Systems
Some AI only gives answers. Other AI can take actions.
AI may edit files, run commands, update records, call outside services, change data, or deploy software. Each step can rely on the result of the previous step.
Anthropic notes that when AI uses connected tools through a series of steps and changes files or systems, mistakes can spread and grow.
Public reports show what this can look like on a smaller scale.
In one Claude Code issue, AI reportedly edited a file that the website page did not use, claimed the work was complete, repeated the mistake, and later ran destructive Git commands while trying to recover.
Another user reported that an AI-generated script deleted fifty audio files that the person intended to review manually.
| Perspective | How It Can Appear | How to Reduce the Risk |
|---|---|---|
| Everyday AI Use | AI sends, schedules, purchases, or updates something before the person reviews it. | Require confirmation before actions that affect another person, money, or public content. |
| Professional Work | One unverified AI result becomes the starting point for later records, reports, or communications. | Check each important result before the next action relies on it. |
| AI-Assisted Technical Work | AI changes shared styles, production data, permissions, dependencies, or deployment settings. | Use changes that can be reversed, limit AI’s access, review exactly what changed, and require approval before any production action. |
AI Execution Risk can produce the following AI failure modes:
- Wrong Tool or Action Selection
- Compounding Multi-Step Errors
- Plans That Don’t Update on New Information
- Security Access & Permission Deviation
- Autonomous Scope Creep
- Insecure Output Handling
- Over-Automation
- Unauthorized External System Access or Action
- Security Safeguard, Monitoring, or Evidence Interference
Human Oversight and Automation Bias: Helpful AI Can Earn Too Much Trust
Useful AI earns trust quickly. That trust can quietly change how people review its work.
After enough good answers, a person may begin checking whether the writing sounds right instead of whether the facts are right.
A neat source list may receive less attention than a messy one.
A passed test may become a substitute for looking at what actually changed.
In 2025, Australia’s Department of Employment and Workplace Relations required Deloitte to correct an assurance report after errors were found in citations and a court-case summary. Government correspondence said generative AI likely contributed to the errors. Deloitte then completed a review directed by its Chief Risk Officer and said the corrections did not change the report’s findings or recommendations. The correspondence also recorded that Deloitte’s policy for telling the client about AI use had not been followed.
| Perspective | How It Can Appear | How to Reduce the Risk |
|---|---|---|
| Everyday AI Use | A person accepts a rewritten message because it sounds better than the original. | Verify promises, tone, facts, and implied meaning before sending. |
| Professional Work | The same AI creates the work and then declares that its own work is accurate. | Ask an independent person, or a separately configured review process starting from the original evidence and acceptance criteria, to assess the work without relying on the first AI’s conclusion. A different AI model may support that review, but it does not replace authoritative evidence or accountable Human judgment. |
| AI-Assisted Technical Work | The reviewer checks the requested page but misses damage to shared files or other pages. | Inspect the actual changes and test every area that could inherit them. Always conduct end-to-end regression testing. |
Human Oversight and Automation Bias can produce the following AI failure modes:
- Skipped or Weak Review Gates
- Over-Trust in Fluent Output
- Unclear Accountability for Sign-Off
Silent Failure Usually Has a Chain
A visible mistake may begin with one AI deviation driver and grow through several others.
A vague request leaves a gap. AI fills it with a likely assumption.
A long conversation weakens an earlier instruction.
The answer arrives with polished confidence. A busy person approves it. AI carries it into another system. Later work is built on top of it.
By the time the customer calls, the original mistake is several steps away.
That is why “AI made a mistake” is rarely a complete explanation.
For professionals designing, building, or governing AI-enabled systems:
When AI makes a mistake, investigate the complete chain of deviations end to end: what the system was expected to do, what context and requirements it received, what assumptions shaped the result, which sources and tools it used, what actions followed, who reviewed or approved them, and how far the effects spread.
Failure mode that may appear later in the chain: Incident Containment, Recovery, or Re-entry Failure
Review AI-Assisted Work Based on Its Possible Impact
Start With Higher-Impact Work
- Customer-Facing Commitments and Policy Statements
- Legal, Financial, Regulatory, Safety, or Employment Content
- Executive Recommendations and Investor Materials
- Requirements, Architecture Decisions, Formulas, and Forecasts
- Published Claims, Statistics, Citations, and Quotations
- AI-Generated or AI-Modified Code, Configurations, Automations, and Data
Use a Fresh Review
Return to the current official source behind every important fact.
Compare the finished work with the approved requirement or intended result.
Identify any facts, broader scope, promises, or certainty that AI added.
Inspect what changed in connected files, records, pages, and systems.
Ask someone who did not create the work, or a different AI starting from the original sources and requirements, to review it without relying on the first AI’s conclusion.
Correct the affected work, then have the correction retested by someone or an AI that did not make the repair.
Record what happened, what else was affected, and which working practice should change.
Do not take shortcuts when checking work that was already completed with AI.
A polished answer is a reason to read carefully, not a reason to skip the evidence.
Expecting Deviations Changes the Question
The useful question is not, “Can AI make a mistake?” It can. It does.
The better questions are:
- Where could a mistake enter this work?
- How far could it spread?
- What would reveal it early?
- Who checks the work before others rely on it?
- If it happens, can we explain it, repair it, and make a repeat less likely?
Expecting deviations does not mean distrusting every AI answer.
It means recognizing that confidence is not evidence, speed is not proof, and correcting the visible output may not correct the weakness that produced it.
The practical standard is simple: Always check important work against trusted sources, involve the right people before decisions are made, and ask what should change so the same mistake is less likely to happen again. That may mean clearer instructions, better information, an added review step, or a safer recovery plan. Fixing the immediate error solves today’s problem. Learning from it helps prevent tomorrow’s.
That is the operating challenge FlyWheel RCA is designed to address.
Sources
Associated Press: AI-Generated Summer Reading List, May 21, 2025
Chicago Sun-Times: Review of the Special Summer Section, May 29, 2025
Meta Instagram Account-Recovery Incident: Reuters, June 3, 2026; TechCrunch, June 1, 2026; The Verge, June 8, 2026; and Meta’s support-assistant announcement, March 2026
The Guardian: PocketOS Production-Database Incident
OpenAI: Why Language Models Hallucinate
NIST: Generative Artificial Intelligence Profile
OpenAI: Expanding on What We Missed With Sycophancy
Anthropic: Effective Context Engineering for AI Agents
IIHS: Study Record on Pedestrian Crash-Prevention Systems
IIHS: High-Visibility Clothing May Thwart Pedestrian Crash-Prevention Sensors
404 Media: Expert-Witness ChatGPT Prompts
Anthropic: Demystifying Evals for AI Agents
GitHub User Report: Claude Code Issue 18891
GitHub User Report: Claude Code Issue 30988
Australian Department of Employment and Workplace Relations: Assurance-Review Correspondence
