Most AI comparisons measure the wrong thing. They line up benchmark scores, compare context windows and crown a winner, and in doing so they answer not a single question that actually comes up in a real company. In this post I put on the table the framework I use to assess AI tools for my clients. Seven criteria, not one of them a benchmark. Every comparison that follows on this blog uses exactly this framework.
The short version
- Benchmarks say little about the value a tool delivers in day-to-day operation, they measure laboratory conditions, not your working week.
- Seven criteria decide it in practice: real workload, context access, data protection, how well it can be rolled out, total cost, exit cost, and the pace of development.
- The license is the smallest cost block: and still the most common anchor for the decision.
- Almost nobody checks exit cost, even though over the years it gets more expensive than the entry price.
- The framework does not make the decision. It only makes sure you ask the right questions before you do.
Why I use a framework of my own
When a new model appears, the comparison tables are online within hours. They show scores on standardized tests: math, programming, reading comprehension. That is not worthless, but it answers a question you never asked. You do not want to know which model does better on a set of exercises. You want to know whether your bookkeeping gets finished faster, and whether the law firm next door will suddenly find your draft contracts sitting in somebody else’s training data.
On top of that: the gap between the leading models has narrowed considerably over the past two years. For the vast majority of tasks in a mid-sized company (emails, summaries, analyses, drafts) every major vendor delivers usable results. The difference that makes or breaks your project lies somewhere else.
A disclosure up front: I earn my living with consulting and training around Claude. That naturally colors my assessments. Which is precisely why this framework is out in the open, you can run the numbers yourself and reach a different conclusion than I do. In the comparisons that follow I also name explicitly, in every post, the cases where another tool is the better choice.
Criterion 1: real workload, not a demo prompt
Every tool looks good in a demo. The question is what happens when you feed it your actual work: the 40-page proposal, the spreadsheet that grew organically over years, the seven-month email thread with three different contacts.
How I test this: I take three to five tasks that come up weekly in the business and have every candidate work through them, with real documents, not sample data. What gets scored is not which result sounds nicest, but how much rework is needed before it is usable. A tool that is brilliant in 80 percent of cases and produces unusable nonsense in the other 20, nonsense you first have to go looking for, is worse in daily use than one that stays consistently solid.
Criterion 2: context access
A language model with no access to your information is a very well-read intern on their first day. It knows the world, but not your company. The real jump in value happens where the tool reaches your data: your files, the mailbox, the document store, the line-of-business application.
The questions that matter: can I store files and folders permanently, or do I have to upload them again in every conversation? Are there connectors to the systems we already run? And can we build our own connection when nothing off the shelf exists for our specialist system, through the Model Context Protocol, for instance?
Criterion 3: data protection and contract terms
This is where the most expensive misconception I meet in consulting sits, and it is not about which vendor you pick, but which plan.
The same pattern holds across the major vendors: on personal plans, content is used to train the models by default, with an opt-out available. On business plans, no training happens on your content by default. In plain terms: the employee who quickly summarizes a client contract using their private free account is the data protection problem, not the vendor choice made at management level.
What I check specifically: is training done on our data, and from which plan onward is it not? Is there a data processing agreement, and who is the contracting party? Where is the data held? How long is it retained, and can that be configured? Which subprocessors are involved?
This is a practitioner’s assessment from consulting work, not legal advice. For a binding evaluation of your specific case you need your data protection officer or a law firm.
Criterion 4: how well it can be rolled out
A tool only the two technically minded colleagues can operate changes nothing on the bottom line. The question is not how good the software is at its best, but how many people will actually use it day to day.
That includes unglamorous things: is there central user management? Can the whole thing be tied to our existing sign-in? Can I, as an administrator, see who uses it, and who does not? And realistically, how much training does it take before someone with no prior knowledge becomes productive? My experience: two to three supported hours per person do more than any feature list. How a rollout like that is structured, I described in the six-week plan.
Criterion 5: total cost over twelve months
The license fee is the item everyone talks about, and the smallest of the three. Do the math like this instead:
- License: price per seat × number of people × twelve months. Watch for minimum quantities and for whether a base license is a prerequisite, with some vendors the advertised price is only the add-on.
- Rollout: selection, setup, policy, connecting it to your systems. One-off, but rarely small.
- Training and support: not the one-time demo, but the support through the first few months, when the questions actually start coming.
For ten people, license costs land in the low four figures per year depending on the vendor. Rollout and training, in my experience, come to more than that. Compare only the license and you are optimizing the smallest lever.
Criterion 6: exit cost
The criterion that almost never appears in a selection process, and that gets more expensive every month. The question is: what does it cost to switch two years from now?
Concretely: can I get my conversation histories, stored documents and self-built templates back out, and in what format? How much work sits in configuration and connections that would be lost in a switch? And how deeply has the tool embedded itself in workflows that would then have to be rewritten?
My rule of thumb: everything you write yourself, meaning policies, templates and curated company knowledge, should sit in a format you can still read without the vendor. Plain text files in your own document store are not a step backwards, they are insurance.
Criterion 7: pace of product development
A double-edged criterion. A high pace means your tool will do considerably more in a year than it does today, without you having to lift a finger. It also means your training material goes stale, that interfaces shift around, and that a feature you built a process on can disappear again.
What I assess here is less the speed than the reliability: does the vendor announce changes? Are there transition periods? Do the APIs stay stable even when the interface changes? For a business that rests its processes on a tool, predictability is worth more than the next feature.
How to apply the framework
Not all seven criteria carry equal weight, that depends on your situation. A tax advisory practice weights data protection and exit cost heavily. An advertising agency weights real workload and context access heavily and can be more pragmatic about data protection, as long as no client data is in play.
My suggestion: assign the weights before you start comparing, which three criteria matter most to you? Only then do you look at the tools. The other way round, an impressive detail regularly ends up shifting the weighting after the fact.
Frequently asked questions
Are benchmarks completely worthless?
No. They reliably show whether a model falls badly short in a category. But for choosing between the three or four leading vendors they are of little use, because the differences there are small and the test conditions are a long way from your working week.
How long does an assessment using this framework take?
For two candidates, realistically two to three weeks: one week for the test tasks, one for the data protection and contract review, plus the time your departments need to give feedback. Decide it in an afternoon and you are deciding on gut feeling, which is sometimes enough, but then please do it knowingly.
Can we not simply use several tools side by side?
You can, it just costs more than most people expect. Not in license fees, but in coordination, training and data protection effort for every additional tool. My recommendation is one standard tool plus a deliberately permitted second one.
Does the framework apply to small businesses too?
Yes, just with less effort. With three people you check the same seven points, but in a few hours rather than weeks. The criteria do not change, only the depth of the review.
So which tool do you actually recommend?
That depends on your weighting, which is what this framework is for. In the following posts in this series I work through the common candidates one by one, each time with the cases where they win and the cases where they lose.
