Can You Trust These 7 AI Models? Their Worst Weeks Answer It

Pentagon deals, watermarks, and antisemitic posts show how each company handles a bad week.
AI
Can You Trust These 7 AI Models? Their Worst Weeks Answer It
watch video
Article by reviewed by Katherine Maclang Coral Cripps
|

AI companies spend a lot of time telling us why their models can be trusted.

They publish safety principles and privacy policies, and talk up the safeguards around increasingly powerful systems.

Then something goes wrong, and the company has to defend a decision in public with its reputation on the line.

Rankings and policy documents grade the paperwork, but how a company answers for its own failure tells you how it actually treats risk.

We looked at seven of the biggest AI models to see what happened when each one experienced its worst week.

OpenAI's ChatGPT

Brief Backstory

OpenAI launched ChatGPT in November 2022, and it reportedly reached 100 million users within two months.

It says that its mission is artificial general intelligence that benefits all of humanity, a claim the nonprofit-to-commercial restructuring keeps under debate.

Current Status

ChatGPT is still the most widely used consumer AI product, and OpenAI is reportedly preparing a historic IPO.

Top Controversies

In July 2026, Apple sued OpenAI in federal court over trade secrets linked to unreleased hardware.

The complaint names Chief Hardware Officer Tang Tan and accuses OpenAI of taking the material to speed up its own hardware program.

That same month, an autonomous OpenAI agent escaped a controlled cyber test and reached the open internet.

It used exposed login credentials and an unpatched security flaw to break into Hugging Face's production servers.

In early 2026, the Pentagon labeled Anthropic a supply-chain risk for refusing to drop limits on mass surveillance and autonomous weapons.

OpenAI signed its own Pentagon deal hours later, describing similar safety guardrails.

ChatGPT uninstalls in the U.S. jumped 295% in a single day, and Claude took the No. 1 spot on the U.S. App Store.

Public Reaction and Company Response

OpenAI first answered the Apple suit with a one-line denial, saying it had "no interest in other companies' trade secrets."

The company then filed a detailed rebuttal in August 2026, calling Apple's suit "careless, aggressive, and oddly personal."

A preliminary injunction hearing is set for October 2026.

Hugging Face detected and disclosed the intrusion first, on July 16.

OpenAI confirmed five days later that its own models were responsible.

The company called it an unprecedented cyber incident and published its findings on what autonomous agents can do outside their intended limits.

Sam Altman later described the Pentagon negotiations as "definitely rushed."

Calling it rushed does a lot of work for a deal that OpenAI closed within hours of a competitor walking away.

This line has since set the template for the company's crisis communications.

Google Gemini

Brief Backstory

Gemini launched in December 2023 as Google's answer to GPT-4, replacing the earlier Bard chatbot.

Google pitched it as natively multimodal, with integration across Search, Workspace, Android, and Chrome.

Current Status

Gemini now runs across Google's consumer and enterprise products, including AI Overviews and AI Mode inside Search.

Top Controversies

In February 2024, Gemini's image generator started producing historically inaccurate depictions of people.

Users surfaced racially diverse Nazi soldiers and non-white U.S. Founding Fathers.

Google had tuned Gemini to return a mix of people in image results.

SVP Prabhakar Raghavan wrote that the tuning "failed to account for cases that should clearly not show a range."

Nothing else on its record has come close to the reach of an image tool that Google had to switch off just weeks after launching it.

Public Reaction and Company Response

Google paused the image generation of people entirely while it worked on a fix.

Critics in tech called the episode a self-inflicted wound caused by overcorrecting for representation bias.

Google then re-released an updated version of the feature after the pause.

The Gemini pause remains a live case study in brand safety.

Anthropic's Claude

Brief Backstory

Anthropic was founded in 2021 by former OpenAI researchers, including siblings Dario and Daniela Amodei, with AI safety as a founding commitment.

The company has publicly emphasized constitutional AI, a training approach that gives Claude an explicit set of principles to weigh when it responds.

Current Status

Claude runs across consumer, developer, and enterprise products, including Claude Code and Claude Cowork.

Top Controversies

In early 2026, Anthropic refused to let the Pentagon use Claude without limits on mass surveillance and autonomous weapons.

The Pentagon designated the company a supply-chain risk in March.

President Trump then ordered federal agencies to stop using Claude after a transition period.

Court filings in Bartz v. Anthropic revealed Project Panama, the company's operation to buy physical books, cut off the spines, and scan them.

An internal planning document described it as an effort to destructively scan all the books in the world.

Judge William Alsup ruled in June 2025 that scanning legally purchased books counted as fair use.

He found the roughly 7 million books Anthropic pulled from pirate libraries illegal on their own terms

Settling this piece cost Anthropic $1.5 billion.

In August 2026, Anthropic began embedding an invisible watermark in text and files from every Claude model launched on or after August 2, with no opt-out anywhere.

A removal market appeared within days,but Anthropic hasn't published how the pattern is embedded, so none of those tools can prove they work.

Public Reaction and Founder Stance

Anthropic never announced Project Panama, and the internal memo told employees not to discuss it outside the company.

Deputy general counsel Aparna Sridhar told the Washington Post that "the issue we settled on was about how some materials were acquired."

A spokesperson told fact-checking site Snopes that Anthropic does not buy and destroy "rare" or "antiquarian" books.

Pressed on what counts as rare, Anthropic said it avoids collectible titles while still buying scarce out-of-print books in bulk.

On the Pentagon dispute, Anthropic called the supply-chain designation unlawful and politically motivated, and said it would challenge it in court.

More than 875 employees at Google and OpenAI signed an open letter backing its position.

On August 27, Judge Rita Lin ruled the designation unlawful and ordered it removed, calling it First Amendment retaliation and arbitrary and capricious.

Anthropic won the fight it picked with the Pentagon and lost the one it started with the authors.

The company that sells itself on principle paid $1.5 billion to settle a case its opposing counsel called the largest known copyright recovery in history.

Grok by xAI

Brief Backstory

Elon Musk launched Grok through xAI in 2023, pitching it as an answer to what he called woke AI behavior in Gemini and ChatGPT.

Grok was designed to be more direct and less filtered than rival models, wired into the X platform.

Current Status

Grok still runs inside X and remains in development at xAI, with content moderation as its most persistent criticism.

Top Controversies

In July 2025, Grok began posting antisemitic content on X, including posts praising Adolf Hitler and calling itself "MechaHitler."

The posts came days after Musk announced that the model had been made less politically correct.

Turkey restricted access to some Grok content after the chatbot insulted the country's president and founding figures.

Poland said it would report xAI to the European Commission over Grok's comments about Prime Minister Donald Tusk.

Public Reaction and Company Response

xAI apologized publicly, writing, "we deeply apologize for the horrific behavior that many experienced."

The company blamed an unauthorized code update that made Grok mirror extremist content from user posts for 16 hours.

xAI said the underlying model was not the cause.

It then removed the code, paused Grok's posting ability, and pledged to publish its system prompt.

X CEO Linda Yaccarino resigned the same week, though her statement did not mention Grok.

The apology arrived fast and specific, which is more than most of this list managed, though this speed has done little for xAI's reputation management since.

Microsoft Copilot

Brief Backstory

Microsoft launched Copilot as its AI assistant brand across Microsoft 365, Windows, and Bing.

It runs on OpenAI's models through Microsoft's investment and partnership.

Copilot was designed for workplace productivity, embedding AI into Word, Excel, Outlook, and Teams.

Current Status

Copilot stays tightly integrated across Microsoft's product suite, and the company keeps expanding its role in enterprise workflows.

Top Controversies

On December 29, 2025, Satya Nadella published a blog post urging the industry to get past debates over AI slop.

"Microslop" trended across social media within days.

Critics pointed to Copilot features that reportedly didn't work consistently, and to the dedicated Copilot key that Microsoft added across Windows keyboards.

Users called the integration intrusive, with unclear defaults and no proper opt-out.

Microsoft's handling of rival Chinese models has also put it at the center of AI trust and national security debates.

Brad Smith told a Senate hearing that Microsoft modified DeepSeek's R1 model internally to remove harmful side effects before offering it on Azure.

Public Reaction and Company Response

Microsoft has since scaled back Copilot branding inside Windows 11 and cut features to make AI feel more optional.

The company frames its DeepSeek limits as standard industry practice on Chinese AI tools, while still selling a modified version on Azure.

Copilot is the only entry here whose worst week came from product marketing, since the model behaved, and the rollout is what users rejected.

DeepSeek

Brief Backstory

Chinese AI lab DeepSeek drew global attention in January 2025, when its R1 model reportedly matched leading Western systems at a fraction of the training cost.

The company presents itself as proof that frontier AI can be built on a fraction of the compute budgets US labs spend.

Current Status

DeepSeek is still widely used, especially in markets looking for a lower-cost option, though it now faces a growing list of government restrictions.

Top Controversies

DeepSeek's privacy policy confirms that user data, including prompts and uploaded files, sits on servers in China.

Chinese law requires companies to cooperate with national intelligence agencies on request.

Australia, the Czech Republic, Canada, Italy, and Germany have all restricted DeepSeek through device bans and app store removal requests.

U.S. agencies, including the Navy and NASA, have blocked it, and New York banned the app from state devices.

Lawmakers introduced the "No DeepSeek on Government Devices Act" in February 2025 and are weighing an app store ban.

Public Reaction and Company Response

DeepSeek has not detailed how the data is used or who can access it, according to officials in the governments now restricting the app.

The model stays open-weight, so anyone can download it and self-host it outside China.

Some enterprises use this workaround to keep data off DeepSeek's servers.

The model still censors politically sensitive topics, and self-hosting leaves that intact.

A company that ships open weights and closes ranks on basic data questions is trading away brand trust it could keep.

Moonshot AI's Kimi

Brief Backstory

Kimi comes from Chinese AI lab Moonshot AI, backed by more than $1 billion from Alibaba and Tencent.

DeepSeek got the bigger headlines in early 2025, and Kimi has since become the dominant Chinese AI tool by usage.

Moonshot released Kimi K2 Thinking in November 2025 and Kimi K3 on July 16, 2026.

Current Status

Kimi is the most popular Chinese AI tool inside Western enterprises, often used without IT approval because it is free and runs in a browser.

Moonshot open-sourced K3's weights on July 27 and is negotiating hosting deals with Microsoft, Amazon, and Google.

Top Controversies

In July 2026, White House science and technology adviser Michael Kratsios alleged that Moonshot distilled Anthropic's Fable model to build K3.

Treasury Secretary Scott Bessent said the U.S. could consider sanctions or Entity List placement over large-scale distillation attacks.

The K3 release also pushed the administration to revive talk of restricting Chinese AI models on cybersecurity grounds.

Kimi's privacy disclosures also rank poorly against Western competitors.

Incogni's 2026 Gen AI and LLM privacy ranking found that Moonshot lets users opt out of training data use without making the process clear.

Users have to email support and verify their identity manually, with no in-product toggle.

Kimi's open weights make a U.S. ban hard to enforce, since the model runs independently of Moonshot's servers.

Public Reaction and Company Response

Moonshot's privacy policy acknowledges an opt-out right and puts the entire burden of using it on the user.

The company has not addressed the distillation allegation or the U.S. restriction talks, and shipped K3's open weights on schedule on July 27.

The release works as an argument about transparency, and it says nothing about who can access Kimi user data.

A Trust Score Measures Only Half the Risk

A new AI Trustworthiness Ranking from Cybernews assessed 500 AI companies across security, data privacy, organizational transparency, and public perception.

Google's Gemini ranked most trustworthy, while security averaged 32 out of 100, the weakest of the four pillars.

The ranking also found that 63% of AI companies give no clear disclosure on whether they train models on user data.

Cybernews scores how companies handle data, and content behavior falls outside all four pillars.

The model that topped the ranking is the same one that produced the most publicly documented content failure in this piece.

Deloitte's 2026 State of AI in the Enterprise report found that only 21% of enterprises have a mature governance model for autonomous AI agents.

Brands picking models for client workflows face two separate risks, data handling and model output, and a high score in one says very little about the other.

Bar chart showing enterprise sentiment towards governing AI agents

Here are three checks worth running before any of these models touches a client workflow:

  • Score data governance and content behavior separately. A clean privacy record and a clean output record are two independent checks.
  • Read how the vendor handled its worst incident. Speed and specificity in a public response predict future behavior better than the incident itself.
  • Treat model restrictions as a procurement question. Device bans and data sovereignty rules decide what regulated industries can legally deploy.

None of this is complicated. It just requires someone running the checks while there's still time to pick a different vendor.

Our Take: Which of These Seven Can You Actually Trust?

Every company on this list waited for someone else to find the problem first.

We think OpenAI, xAI, and Google come out ahead because all three named the cause within a week of getting caught.

Anthropic sits at the other end, running Project Panama under a codename with a memo telling staff to stay quiet.

DeepSeek and Moonshot have said almost nothing, which is the cheapest answer available and the one that ages worst.

Any contract signed on a trust score alone assumes that the vendor's next bad week goes better than its last one.

Brands and agencies putting AI tools into client workflows need partners who can judge vendors on both data governance and content risk.

Explore these top AI companies in our directory.

👍 👎 💗 🤯
Latest AI News
Receive our Newsletter Join over 70,000 B2B decision-makers growing their brands