Last month I did what I always do at the end of a quarter: I opened a document called content_ideas_raw.docx and stared at it.
It had 63 ideas in it. Some were full sentences. Some were three words and a question mark. Two were just the word "newsletter??" with no further explanation. I had collected them from meetings, Slack messages to myself, competitor emails, and the notes app on my phone at 11 p.m.
I did not want to write 63 articles. I wanted to know which 8 or 10 were worth pursuing. So I handed the list to AI and asked it to sort a month of content ideas into something usable.
Here is what the tool actually did — and where I still had to do the work myself.

The Task and the Conditions
What I was trying to finish: Turn a messy, 63-item idea dump into a prioritized shortlist for the next month of content, with rough categories and a reason for each ranking.
Tools used:
ChatGPT (Plus plan, GPT-4 class model) — main sorting and clustering
Claude (free tier) — second opinion on the top 15
Google Sheets — where the final list lived
Input: A plain-text list of 63 ideas, copied directly from my notes. No formatting, no tags, no dates.
Time limit I set for myself: 45 minutes total, including prompting and reviewing.
What I did not give it: My company's content calendar, traffic data, or any client names. I stripped out anything specific before pasting. (More on that in the guardrails section.)
The Workflow, Step by Step
Step 1: The first prompt (and why it was too vague)
My first prompt was: "Organize these content ideas into categories and tell me which ones are best."
The output looked impressive. It gave me six tidy categories with names like "Educational," "Thought Leadership," and "Industry Trends." It also told me the "best" ideas were the ones with the most words in them.
That was useless. The longest idea in my list was a rambling two-sentence note I had written at a conference and never acted on. Length is not quality. I had asked a sorting question and gotten a sorting answer — but not a useful one.
Lesson one: "Best" is not a brief. I had to tell the tool what "best" meant for me.
Step 2: The second prompt — adding real criteria
I rewrote the prompt to include three things:
Audience: marketing and content professionals at small and mid-sized companies
Goal: ideas that could become a practical, first-person article with a testable angle
Constraints: avoid anything that required original data I did not have, and avoid pure news commentary
I also asked for output in a table with four columns: Idea, Category, Why it fits, What is missing.
This was better. The tool grouped the 63 ideas into five clusters that actually made sense:
Workflow experiments (testing a tool or process)
How-to explainers (teaching a specific skill)
Opinion / debrief (lessons learned)
Reactive / news-adjacent (I ended up cutting most of these)
Unclear (ideas that were too vague to place)
That fifth category was the most honest thing the tool did all day. It did not pretend to understand "newsletter??" — it put it in Unclear and moved on.
Step 3: Asking for a shortlist with reasons
Next I asked it to pick the top 12 from the clusters, and for each one, to tell me:
What the article would actually be about in one sentence
What the reader would get
What I would need to test or verify before writing it
This is where the tool started earning its keep. It surfaced three ideas I had buried near the bottom of my list and had completely forgotten about. One of them — a note about testing whether AI could clean up meeting notes — became the basis for an article I am now planning.
But it also did something I did not ask for.
What Went Wrong
It ranked by "sounds good," not by "fits my actual situation"
The tool's top pick was an idea about AI and team productivity. It sounded strong on paper. But I do not manage a team, I do not have access to team productivity data, and I could not write that article without either inventing details or exposing my workplace. It was a good idea for a different writer.
The tool had no way to know that. I had given it the ideas but not my constraints as a writer. That was my omission, not the tool's failure — but it is exactly the kind of gap that makes AI output look better than it is.
It flattened distinct ideas into the same category
Two of my ideas were genuinely different: one was a tool review, and one was a workflow reflection. The tool put both under "AI Productivity" and treated them as interchangeable. When I asked it to expand the shortlist, it essentially wrote the same article concept twice with different titles.
I caught this because I knew my own list. A reader who trusted the output without checking would not have.
It could not tell the difference between an idea and a note
Some entries in my list were not ideas at all. "Ask about the new dashboard" was a task, not an article. "That thing Sarah said about tone" was a reminder, not a topic. The tool tried to turn both into content angles. One of them it turned into a genuinely weird article pitch about "dashboard-driven content strategy."
That one made me laugh, and then it made me check every single item by hand.

What Still Needed My Attention
After about 40 minutes of prompting and reviewing, I had a shortlist of 14 ideas. Here is what the human editing actually consisted of:
What AI did | What I still had to do |
|---|---|
Grouped 63 ideas into 5 clusters | Delete the ones that were tasks, not ideas |
Wrote one-sentence summaries | Rewrite 6 of them because they missed the point |
Ranked by apparent strength | Re-rank based on what I could actually test and publish |
Flagged missing information | Decide which gaps were fatal and which were fine |
Suggested categories | Rename two categories to match my site's actual structure |
The time estimate: roughly 25 minutes of AI work, 20 minutes of my own review. I did not save hours. I saved the worst part — staring at a blank page trying to figure out where to start.
The Verdict
Keep it, use it selectively.
I will use AI again for this task, but only for the first pass: clustering and surfacing forgotten ideas. I will not use it to rank or decide. The ranking step is where the tool's lack of context does the most damage, and it is also the step where my judgment is cheapest to apply — I already know my constraints.
What it is good at: turning a messy pile into visible groups, and reminding me of ideas I forgot I had.
What it is not good at: knowing which ideas are actually writable, which ones fit my voice, and which ones are just notes to myself.
The limitation to remember: AI sorts what you give it. It does not know what you can realistically do with the result. If your input is messy in a way you have not named, the output will be confidently messy in a way that looks organized.
The Anonymized Version of This Workflow
If you want to try this yourself, here is the prompt structure that worked after the first one failed:
"Here is a list of [number] content ideas. My audience is [describe them]. My goal is [what you want the reader to get]. I can write about [what you have access to] but not [what you do not]. Please group these into categories, flag any that are too vague to use, and give me a shortlist of [number] with a one-sentence summary and a note on what I would need to verify before writing."
The key line is the one about what you cannot write about. That is the constraint the tool will not guess, and it is the one that saves you the most time.
Guardrails I Followed
I stripped all company, client, and coworker names before pasting anything into a tool. The input was just the idea text.
I did not use any real internal project names, even in anonymized form.
I treated every ranking as a suggestion, not a decision.
I checked each shortlisted idea against what I could actually test and publish before committing to it.
The first draft of the shortlist came from AI. The judgment did not.
Next in this series: I gave AI my presentation outline before I gave it my trust. Here is what came back.
I tried it so you don't have to waste your afternoon.
No notes yet — be the first to inscribe one.