July 20, 2026 · Vannus · ~5 min read · All posts

We added 69 AI tools, and held them back twice before publishing.

Two attempts would have published failing grades on real companies because of gaps in our own database. Here is what went wrong, what we changed in the methodology, and why every tool page now carries two verdicts instead of one.

The catalog went from 298 tools to 367 today. That sentence took three attempts, and the two failures are more useful than the success.

The first attempt would have published 67 F grades

We researched 69 tools that companies are actually running in 2026, weighted deliberately toward vendors outside the United States. A catalog about who controls your AI stack is worth very little if every alternative it lists answers to the same jurisdiction.

The research went well. Forty-five of the sixty-nine carried a country of origin taken from the vendor's own site with a source URL recorded; not one country claim was inferred from a domain suffix or a founder's name. Six vendors that train on customer data by default were left with no training value at all, because our vocabulary has no term for it and the fallback in our scoring code is the favourable reading — an unset field says "not assessed", which is true, where a default would have been a flattering falsehood.

Then we scored them.

 Existing 298New 69
F grades1567
F rate5.0%97.1%

Ninety-seven percent. Not because those vendors are bad — because we had not recorded their pricing or their integrations, and the scoring engine read a missing field as a bad answer. Measured on one entry under the engine as it then was: Google Jules scored 11 out of 100 with those fields absent and 39 with them present. The same product, the same day, twenty-eight points of difference sitting entirely in our database rather than in their software. Those two figures are a record of the defect, not of the current engine — after the fix described below the same comparison reads 20 and 44, and Google Jules publishes at 44.

We publish letter grades on named companies. Publishing an F because our own row was empty is not a rating, it is a defamation risk with a letter on it. We stopped and reverted.

What we changed in the methodology

Our stated rule has always been that where a vendor discloses nothing, we say so rather than guessing. That rule was enforced for one criterion — which model a product runs on — and quietly ignored everywhere else. Three changes fixed that.

An axis with no data is excluded, not scored zero. Two criteria were scored unconditionally, so a tool whose integrations and pricing we had not recorded lost 35 of 100 points before anyone looked at it. They now behave like the model-provenance criterion already did: absent data drops out and the remaining criteria renormalise. Across the live catalog this moved 7 grades of 298, and every one moved up. No vendor's published grade got worse.

Below half the framework, we publish no letter at all. Excluding empty criteria was not enough on its own — a bare entry still scored 23 on what remained, which is an F for the offence of being newly added. A grade implies we looked. Under 50% coverage the honest statement is that we have not finished looking, so those tools now show "not yet graded" and no score. One of the 69 is in that state today. It is listed, searchable, and ungraded, which is the truth.

We publish how much of the framework each grade rests on. This one is a confession. Renormalising over fewer criteria raises the average whenever the dropped criterion was below it. Writer scores 79 fully assessed and 82 with two criteria stripped out — so a thinly researched tool can grade slightly higher than a thoroughly researched one. That is the real cost of refusing to score absent data as zero, and the alternative is worse. We are not going to solve it by pretending, so every grade now shows the number of criteria it was assessed on. If it says 4 of 6, weigh it accordingly.

The second failure was more interesting

With the research completed and the methodology fixed, the F rate fell from 97.1% to 11.6%. We were ready to publish. Then we read the results.

 Proton LumoGlean
CountrySwitzerlandUnited States
Within U.S. CLOUD Act reachNoYes
Trains on your dataNeverNever
Integrations110
GradeFA+

Both grades are arithmetically correct. Our letter grade measures resilience — whether a tool endures, whether its pricing is durable, whether you could leave it, how deeply it connects to the rest of your stack. Proton Lumo has one integration by design, because that is what a privacy-first product looks like. It earns a low resilience score honestly.

But that F was about to appear on a page headed "Who really controls your AI stack?", on a site whose paid product is a sovereignty report. A reader arriving with a sovereignty question would have read a resilience answer and concluded that the Swiss, non-US, never-trains option was the worst on the list. That is the opposite of true, it is unfair to Proton, and it would have been our fault, not theirs.

We checked whether the engine simply favours American vendors. It does not: non-US-controlled tools average 66.0 across the existing catalog against 58.1 for US-controlled ones. The defect was not bias. It was that we were answering a question nobody asked us.

So every tool page now carries two verdicts

The letter grade stays exactly what it is, and keeps its name: a resilience grade. Next to it, every page now states the sovereignty position in its own words — country of origin, data jurisdiction, whether the vendor sits within reach of the U.S. CLOUD Act, and whether it trains on your data.

Not a second letter. Two letters side by side would have moved the confusion rather than removed it. A label answering a different question, with a line on the page saying plainly that it is a different question: the grade measures whether a tool endures and whether you could leave it; the sovereignty panel measures who can compel your data. A tool can score modestly on one and strongly on the other, and a great many do.

Where a vendor documents nothing, the panel says "not yet assessed" — and says that we do not infer sovereignty from a domain name or a company name, because we do not.

Why write this down

Because the failures are the product. Anyone can publish 367 grades. The question a buyer should ask is what happens when the grades are wrong, and the only useful answer is a record of what we did the last time they were.

Twice this week the fastest path was to ship and see if anyone noticed. Both times the thing standing in the way was a number that looked wrong — a 97% failure rate, and an F on the most privacy-respecting vendor in the batch. Neither was caught by a customer complaint, because we have not got that far yet. They were caught by looking.

If you think a grade of ours is wrong, there is now a page for that: corrections and right of reply. It is free, it does not need a lawyer, and if you disagree with our judgement rather than our facts we will publish your response next to our grade, unedited. Stale bad news is the correction we most want to hear about.

— Drake Schreiber, founder, Vannus (PRAXIS AI LLC)