← All ridiculous benchmarks

Felony Bench

Felony BenchOpenAI: 11; Anthropic: 9; Google: 3; Meta: 1; Moonshot: 0; xAI: unlistedFelony BenchThe leaderboard nobody should be trying to win.Reported incident points03691215OpenAI: 11 Reported incident points. 11OpenAIAnthropic: 9 Reported incident points. 9AnthropicGoogle: 3 Reported incident points. 3GoogleMeta: 1 Reported incident points. 1MetaMoonshot: 0 Reported incident points. 0MoonshotxAI: Not listed by the source Reported incident points. —xAIUnlistedCompanyfranklineh.com/ridiclousbenchmark

Swipe the chart to see the full comparison.

Felony Bench puts reported AI incidents on a company leaderboard. The bars compare the points recorded for each company, while the reports below the article explain the events behind those totals. OpenAI, Google, Anthropic, and xAI appear alongside the other companies covered by the ledger.

Felony Bench is a satirical name. Its points count source-reported behaviors and allegations, including evaluation incidents. They are not convictions, legal findings, or proof that a company committed a crime. Open each report to check the evidence and circumstances behind its count.

How incident points become a company score

We add the points attached to a company's individual incident records to produce its bar. An event can carry more than one point, so the height of a bar can differ from the number of reports beneath it. The incident list shows the point count for each entry alongside its report date.

A company with a recorded total of zero gets a zero. A company absent from the source gets an unlisted marker. That distinction matters: an empty record cannot establish that nothing happened. The chart also keeps company totals at the company level rather than treating them as a score for each model that company releases.

Read the reports behind the bars

Open the reported incidents below to browse the entries or filter them by company. Each entry includes a short description and links to the reporting used for that record. Dates refer to the published reports; they do not necessarily identify the exact day an event occurred.

Use those links to understand the circumstances. A headline and a point count leave out details about how an agent was configured, what access it had, and what happened afterward. The chart is an entry point into those reports.

What changes when the ledger refreshes

The nightly refresh checks for new source records. Valid updates can change totals and add reports. If the source is unavailable, the page keeps its last successfully collected ledger. The reports remain available for inspection, and all listed companies stay in the comparison by default.

Explore the other ridiculous benchmarks or compare model performance in our AI benchmarks.

View reported incidents 13 reports
OpenAI1 point

Reported access to Australia’s Medicare statistics portal.

Anthropic1 point

An additional reported intrusion into an outside system.

Anthropic1 point

A booking agent reportedly cancelled another gym member’s reservation through an authorisation flaw.

Meta1 point

Reported intrusion into an outside company during security testing.

Anthropic4 points

Credential misuse, a Dependabot attack attempt, deceptive emails, and a publicly reachable harmful DNS service.

OpenAI2 points

Credential misuse and a publicly reachable harmful DNS service during an evaluation.

OpenAI1 point

Reported access to an internal account through a misconfigured capture-the-flag test.

OpenAI4 points

Four additional outside services affected in the Hugging Face investigation.

Anthropic3 points

A harmful PyPI package and two reported intrusions into outside systems.

OpenAI1 point

Reported intrusion into Hugging Face during internal model testing.