I have cancelled more AI subscriptions than I have kept. Not because the tools were bad, but because most of them were fine in a way that did not justify a monthly charge.
The problem with how most people decide is that they evaluate tools in the wrong order. They sign up, poke around, feel impressed, and then try to figure out whether it fits their workflow. By the time they realize it does not, they have paid for two or three months and built a small habit around something they do not need.
I do it the other way now. I run a short test before I decide anything, and the test has three possible outcomes: buy it, try it, or skip it. Here is the whole thing, in the order I actually run it.

The Test, at a Glance
Before I open the tool, I answer one question: what specific task am I testing this on?
Not a category. A task. "Drafting weekly reports" is a category. "Turn these eight bullets into a first draft I can edit in twenty minutes" is a task. If I cannot name the task, I am not ready to test anything, and I skip the tool until I can.
Once I have the task, I run four checks:
The 30-Minute Real Task — does it help on something I actually have to do today?
The Check Cost — how much time do I spend verifying the output?
The Day-Three Test — do I still reach for it when the novelty is gone?
The Cancel Question — if I had to pay for this out of my own pocket, would I?
The first two decide try it or skip it. The last two decide buy it.
Check 1: The 30-Minute Real Task
I take one task I have to do anyway, set a 30-minute timer, and use the tool for it. No demo prompts. No best-case inputs. The actual messy thing on my desk.
This matters because tools are usually tested on clean inputs. A real task has a half-finished document, an unclear goal, and a deadline. If the tool only shines on clean inputs, it will not shine in my actual week.
What I am looking for: did I finish the task faster, or did I finish it at all when I otherwise would have stared at it?
What disqualifies a tool: if I spend the 30 minutes learning the interface instead of doing the task, that is not a failure of the tool. It is a sign I have not tested it yet. I extend once, and only once.
Check 2: The Check Cost
This is the check most people skip, and it is the one that decides everything.
Every tool has a checking cost. If AI drafts something, I have to read it, verify the facts, fix the tone, and confirm it says what I meant. That cost is real, and it is usually invisible until you measure it.
My rule: if the checking cost is higher than the drafting cost, the tool is not helping. It might be producing output, but I am still doing the work, plus a review pass.
I measure this crudely. I note roughly how long the draft took to produce, and roughly how long it took to check and fix. If the second number is close to or larger than the first, I do not count the tool as a time-saver for that task. I might still use it, but for a different reason — usually getting past a blank page, not saving time.
What disqualifies a tool: output that looks right and is wrong in ways I only catch on a second read. That is the worst case, because the checking cost is high and the risk is hidden.

Check 3: The Day-Three Test
Novelty is a terrible judge of tools. Almost everything feels useful on day one, when you are exploring and the interface is new.
So I wait. I use the tool on real tasks for three days. Not three sessions — three days, with at least one day where I do not open it at all.
On day three, I ask one question: did I reach for it out of habit, or did I have to remind myself it existed?
If I had to remind myself, that is data. It usually means the tool solves a problem I do not have often enough to build a habit around.
What disqualifies a tool: if the only reason I opened it on day three was to complete the test, the test is already over. It is a skip.
Check 4: The Cancel Question
The last check is the one that has saved me the most money.
I imagine the charge hitting my personal card. Not the company card, not a free trial, not a promotional rate. My card, at full price, every month, starting now.
Then I ask: would I renew this right now, knowing what I know?
This question cuts through a lot of noise. A tool can be useful and still not worth a subscription. A tool can be impressive and still not worth the friction of another login, another bill, another thing to manage. "Useful" and "worth paying for" are different questions, and most tool reviews collapse them.
What disqualifies a tool: if my honest answer is "I would keep it if it stayed free," it is a skip, not a buy.
The Three Verdicts
After the four checks, the tool lands in one of three buckets.
Buy It
The task is recurring, the checking cost is lower than the drafting cost, I reached for it on day three without reminding myself, and I would renew it at full price right now.
This is rare. Most tools I test do not get here. The ones that do tend to be narrow — they do one specific thing I do often, and they do it well enough that I stopped thinking about whether to use them.
Try It
The task is real but not frequent, or the tool helps in some cases and not others, or I need more time to see whether the checking cost holds up.
This is the most common outcome, and it is the one people mishandle. "Try it" does not mean "keep paying while I figure it out." It means I set a decision date — usually two weeks out — and I put the tool on a short leash. If I have not built a habit around it by then, it becomes a skip. No extensions without a specific reason.
Skip It
The checking cost is too high, or I had to remind myself the tool existed, or I would not renew at full price.
Skipping is not a judgment that the tool is bad. Some tools I skip are genuinely impressive. They just do not fit a task I actually have. Writing that down honestly is the whole point.
What Went Wrong When I First Started Testing This Way
I tested tools instead of tasks
Early on, I would open a new tool and explore it, then try to imagine where it fit. That is backwards. It produces a list of features and no answer to whether I will use it. Testing tasks first fixed this.
I confused "free trial" with "low cost"
A free trial still has a cost: the attention it takes to evaluate, the habit it might build, and the cancellation you have to remember. I now treat trials as a real decision, not a free pass.
I kept tools because I liked them, not because I used them
This is the quiet one. A tool can be pleasant to use and still not earn its place. The cancel question is the only check that catches this reliably, because it forces me to imagine actually paying.
I let "try it" become permanent
The try-it bucket is where subscriptions go to survive past their usefulness. Setting a decision date fixed that. No date, no trial.
What Still Needed My Attention
What the test does | What I still had to do |
|---|---|
Names the task before testing | Refuse to test tools that have no task yet |
Measures rough check cost | Be honest when the check cost was higher than the draft cost |
Runs the day-three habit check | Notice when I was only opening it to finish the test |
Forces the cancel question | Say no to tools I liked but would not pay for |
Sorts into three verdicts | Set a real decision date for every "try it" |
The Limitation to Remember
This test tells you whether a tool earns a place in your workflow for a specific task. It does not tell you whether the tool is good in general, and it is not meant to. A tool that fails my test might be perfect for someone with a different task, and a tool that passes mine might fail yours.
That is not a flaw in the test. It is the point. The only question that matters is the task in front of you.
The Anonymized Version of This Test
If you want to run it yourself, here is the short version:
Task: what specific task am I testing this on?
1. 30-minute real task: did it help on something I actually had to do?
2. Check cost: is the checking cost lower than the drafting cost?
3. Day three: did I reach for it out of habit?
4. Cancel question: would I renew this right now at full price?All four yes → buy it. Mixed → try it, with a decision date. Otherwise → skip it.
Guardrails I Followed
I did not paste any employer, client, or coworker details into any tool to write this piece. All examples are generalized.
I did not present this test as a way to judge whether a tool is good in general. It only judges fit for a task.
I did not recommend any tool I have not personally tested.
I did not let a free trial count as evidence of value. The cancel question is the one that decides.
The verdicts came from the test. The test is mine.
This is the twentieth post in the launch roadmap. If you want the full set of experiments behind it, the series pages have them.
I tried it so you don't have to waste your afternoon.
No notes yet — be the first to inscribe one.