July 20, 2026 · Vannus · ~5 min read · All posts

We re-graded every AI tool in our catalog. 233 of 287 grades changed.

A rating system that never revises itself isn't rigorous. It's decorative. So here is a revision, in public, with the error named.

The heaviest criterion in our scoring — the one meant to measure whether a tool is built on its own AI or is a thin reseller of someone else's model — was measuring the wrong thing. It credited a tool based on whether its name appeared on a list. It never read what the tool actually runs.

The result was worse than useless. Measured against the evidence, the criterion was inverted: tools that provably resell a third-party model averaged more credit on that axis than tools that provably build their own. Nearly half the catalog carried a failing grade largely because a name wasn't on a list.

What we changed

The criterion now reads what the vendor itself publishes about the model behind the product — and where the vendor discloses nothing, it says so instead of guessing. We researched all 287 tools against vendor primary sources and re-graded the catalog.

94 → 15
tools graded F
233
grades changed
149
grades cite vendor docs

The reversals

Some of the tools we had marked hardest were the ones we had most wrong. Two examples, with the receipts:

Notion AI is routinely called "just a ChatGPT wrapper." It isn't — and Notion's own security page says so:

various large language models (LLMs) hosted by Notion as well as by organizations such as Anthropic and OpenAI
notion.com/help/notion-ai-security-practices · vendor-sourced

GitHub Copilot was carrying a wrapper-tier grade too. GitHub's own model documentation names five model providers, plus Microsoft's own model. A tool routing across five providers is not a single-upstream reseller. Both tools' grades moved up.

The rule that replaced the old behavior

Here is the part that matters more than any single grade. We now publish a negative-sounding claim about a named company — "this tool depends on one upstream model" — only where we can produce the vendor's own document saying so. Where a vendor doesn't disclose what it runs, we don't guess and we don't imply. We mark it not disclosed and grade it on the criteria we can actually check.

That is the opposite of how most "AI tool" marketing works, and it's the whole point of Vannus. The grade is a conclusion. The vendor's own words are the receipt. You shouldn't have to trust the first without being able to check the second.

Why tell you at all

Because the alternative — quietly changing 233 grades and hoping nobody noticed the old ones — is exactly the posture we grade other tools for. If our scoring can be wrong, the honest move is to show you when it was, and what fixed it. That's the version of a ratings company worth paying attention to.