HumanizerBench
HumanizerBench compares AI rewriting tools using detector results and text-quality measures. The site publishes monthly leaderboards and says it makes prompts, raw outputs, detector verdicts, and scoring scripts available for inspection. WriteHuman, one of the tested products, operates the benchmark.
Read the formula behind the ranking
The displayed September 2026 cycle covers fourteen tools, thirty-three prompts per tool, and five detectors. Its overall score weights detector bypass at 42%, meaning preservation at 32%, readability at 16%, and consistency at 10%, with penalties for flagged output problems. Those weights are choices made by the benchmark operator, not a universal definition of good writing.
Meaning preservation uses embedding similarity; readability uses a language-model rating. A detector verdict answers a different question from whether an article is accurate, well sourced, or appropriate for its intended audience. Do not treat a high overall score as evidence that a rewritten claim remains true.
Inspect the actual rewritten passages
The site lists penalties for meaning drift, length inflation or deflation, unchanged output, and refusal. Examining the examples behind those penalties can be more useful than choosing a tool from its rank alone. Compare the rewrite with the input to see whether a qualification or concrete detail disappeared.
Detector-specific pages and category rankings reuse the underlying cycle data. The site's Turnitin-oriented ranking uses Originality.ai as a proxy; it is not a direct Turnitin test. Franklin has not reproduced the scoring run or verified performance outside the published prompt set.
Keep the operator's interest visible
The about page discloses WriteHuman's role and states that editorial control is separate from its commercial product team. It describes public, deterministic scoring and no paid placement. These are disclosed governance commitments, not proof that conflicts of interest disappear.
Use the public data to assess the comparison rather than relying on an independence label. Monthly results reflect a particular test cycle and can change with tool or detector updates. Editing for clarity should preserve attribution and factual meaning; detector scores do not replace those editorial responsibilities.
