A blog archive gets hard to reason about long before it gets large. You remember the newest posts, but older pages sit in a fog: some are still useful, some need a factual check, and a few cover nearly the same ground.
An AI content audit can help you sort that archive, provided AI does not make the final decisions. The output in this workflow is a simple spreadsheet that records what each page does, the evidence available, and one next action. It is an editorial workbench, not an automatic verdict on which URLs deserve to stay.
You need a list of your own published URLs, a spreadsheet, and an AI assistant that can analyze text you provide. Search analytics are optional. A small blog can complete a useful first pass with page titles, dates, summaries, and manual notes.
Table of Contents
- Decide what this audit should answer
- Build a spreadsheet that separates facts from judgments
- Collect the inventory without creating a data swamp
- Run the seven-step AI content audit
- Prompts for classifying and reviewing pages
- Worked example: four posts from a hypothetical baking blog
- Turn labels into a realistic refresh queue
- Where AI audit suggestions go wrong
- Human review checklist
- Questions bloggers ask about content audits
- Audit a small batch before the whole archive
Decide what this audit should answer
Do not begin by importing every metric you can find. Start with the decision you need to make.
For a small blog, a first audit usually needs to answer five practical questions:
- What is each page mainly about?
- Is the information still accurate and useful?
- Does another page serve nearly the same reader need?
- What work, if any, should happen next?
- Which pages deserve attention first?
This workflow fits bloggers and small editorial teams that can inspect their own pages. It is not a substitute for a technical crawl of a large site, legal review, accessibility testing, or a professional analytics investigation. It also will not predict rankings.
Choose a scope before you open the sheet. You might audit one category, posts older than a year, or the 30 pages with the most impressions. A narrow batch is easier to review, and mistakes in your labels will not spread across the whole archive.
Google’s guidance on helpful, reliable, people-first content includes self-assessment questions about original value, sourcing, expertise, and whether a reader leaves with a satisfying answer. Those questions are useful editorial prompts. They are not a scoring formula, and checking more boxes does not promise better search performance.
Build a spreadsheet that separates facts from judgments
A good audit sheet makes uncertainty visible. Put imported or directly observed facts in one group of columns. Put AI suggestions and editor decisions in another.
| Column | What belongs there | Source | Who approves it |
|---|---|---|---|
| URL | Canonical public page address | Site inventory | Editor verifies |
| Title | Current published title | Page or export | Editor verifies |
| Publish/updated date | Dates shown by the site | Page or CMS | Editor verifies |
| Primary reader task | One sentence describing what the page helps with | Page text | Editor |
| Accuracy flag | Current, needs verification, or known outdated | Source review | Editor |
| Overlap note | Possible relationship to another page | AI comparison plus reading | Editor |
| Suggested action | Keep, refresh, combine, redirect-review, or archive-review | AI suggestion | Editor |
| Priority | Now, next, later, or no action | Evidence and effort | Editor |
| Reason | Short evidence-based explanation | Audit notes | Editor |
| Owner/status | Person and production state | Team | Editor |
Add optional analytics only when they help the decision. Google Search Console’s official Performance report documentation describes clicks, impressions, click-through rate, and average position, along with ways to group results by page and query. These numbers need context. A page with few clicks may answer a narrow but important question; a page with impressions may still contain an outdated fact.
Keep blanks as blanks. Do not turn missing traffic data into zero, and do not ask AI to estimate a metric it has not received.
Collect the inventory without creating a data swamp
Begin with one row per public article in scope. You can copy URLs from your sitemap, a CMS export, or a hand-built list. For a very small site, manual collection is fine.
For each row, capture the title, visible date, category, and a short excerpt or summary. If you plan to analyze full text, work in batches and check your AI tool’s input and privacy rules first. Remove private notes, customer data, unpublished client material, and credentials.
Then add evidence that only a person can verify well:
- broken or obsolete external references;
- product screenshots that no longer match the interface;
- dated advice, prices, policies, or statistics;
- missing author experience where the article implies first-hand knowledge;
- comments or support questions that reveal a confusing section.
Do not force every URL into an action immediately. The inventory stage should describe the archive. Decisions come after you can compare related pages.
If an older article clearly needs work, the Practical AI Flow AI-assisted content refresh workflow shows how to verify claims and revise the page without replacing useful material blindly.
Run the seven-step AI content audit
1. Define the decision rules
Write plain definitions for each action. For example, refresh means the reader need is still valid but facts, examples, structure, or links need work. Combine candidate means two pages appear to serve substantially the same intent, but an editor must read both before deciding. Avoid a blunt “delete” label.
2. Prepare a small evidence packet
Select 10 to 20 rows. Include the URL, title, date, summary, target reader, and any verified notes. Give every row a stable ID so the response can be matched back to the sheet.
3. Ask AI to normalize topics and reader tasks
Have the model write one topic label and one reader-task sentence for each row. This makes inconsistent titles easier to compare. The model must use only supplied text and mark uncertain cases.
4. Look for possible overlap
Ask for pairs or groups that may answer the same question. Require a reason and a confidence label. Similar words alone are weak evidence: “blog outline” and “content brief” may be related without being duplicates.
5. Propose an action, not a command
For each row, request one suggested action and the evidence behind it. Allow “keep as is” and “manual review needed.” An audit becomes distorted when the prompt assumes every page needs changing.
6. Review the riskiest decisions first
Read every combine, redirect, archive, or major-rewrite suggestion. Open the live pages and check their purpose, links, sources, and any traffic context you use. AI cannot know the value of a page from a thin summary.
7. Assign priority and ownership
Rank approved tasks by reader harm, factual urgency, strategic fit, and effort. A wrong policy date can deserve attention before a promising title rewrite. Add an owner and status so the sheet becomes a working queue rather than a report that disappears in a folder.
Prompts for classifying and reviewing pages
Use this first prompt on a small inventory batch. Replace the bracketed fields and preserve row IDs.
You are helping with an editorial content audit. Analyze only the supplied
inventory. Do not invent traffic, conversions, rankings, dates, or page content.
For each row, return:
- Row ID
- Normalized topic (2-6 words)
- Primary reader task (one sentence)
- Accuracy status: no issue visible / verify / known issue
- Suggested action: keep / refresh / possible combine / manual review
- Reason based only on supplied evidence
- Confidence: high / medium / low
- Missing evidence needed before a decision
Use "manual review" when the evidence is too thin. A similar title is not enough
to recommend combining pages.
<inventory>
[PASTE 10-20 LABELED ROWS WITH TITLES, SUMMARIES, DATES, AND EDITOR NOTES]
</inventory>
After checking those row-level labels, use a separate comparison prompt. Keeping the jobs separate makes it easier to spot a bad assumption.
Compare the approved audit rows below for possible content overlap.
Return a table with:
- Row IDs compared
- Shared reader need
- Important difference in scope or audience
- Relationship: distinct / supporting pair / possible overlap
- Evidence from the supplied summaries
- What a human must check on the live pages
Do not recommend deletion, redirection, or merging. Do not use keyword similarity
alone. If two pages belong in the same topic cluster but solve different tasks,
label them "supporting pair."
<approved_rows>
[PASTE REVIEWED ROWS]
</approved_rows>
A final scheduling prompt can help order work after the editor approves each action:
Create a four-week editorial queue from these human-approved audit actions.
Prioritize: (1) known factual risk, (2) broken reader journey, (3) important
missing explanation, then (4) routine polish. Respect the effort estimate and
weekly capacity supplied below. Do not change any approved action.
Return: week, row ID, action, reason for timing, prerequisite, owner placeholder.
Weekly capacity: [HOURS OR NUMBER OF PAGES]
Approved actions: [PASTE ROWS]
Worked example: four posts from a hypothetical baking blog
Consider a fictional beginner baking site with four inventory rows. No real site or private analytics are involved.
| ID | Title | Date | Supplied note |
|---|---|---|---|
| B01 | Sourdough Starter for Beginners | 2024-02 | Clear steps; one dead reference link |
| B02 | Fix a Weak Sourdough Starter | 2025-09 | Troubleshooting guide with symptom table |
| B03 | Easy Sourdough Starter Guide | 2022-03 | Similar opening to B01; old temperature advice not yet verified |
| B04 | Baking Tools for Small Kitchens | 2025-11 | Equipment checklist for apartment bakers |
The representative classification prompt returned a possible-combine suggestion for B01 and B03, kept B02 as a distinct troubleshooting page, and marked B04 as separate. That is a reasonable lead, but the raw result was too confident. It treated title similarity as proof and suggested a redirect before anyone had compared the full pages.
The revised audit entry reads:
| Rows | Editorial decision | Priority | Reason | Next check |
|---|---|---|---|---|
| B01 | Refresh | Next | Core instructions appear useful; replace dead reference | Verify replacement source |
| B02 | Keep | No action | Solves a troubleshooting task rather than initial setup | Recheck during annual audit |
| B03 | Manual overlap review | Now | Possible duplication plus unverified temperature advice | Read B01/B03 fully and verify advice |
| B04 | Keep | Later | Different reader task; no issue in supplied evidence | Check links in routine review |
Notice what changed. The model’s irreversible recommendation became a reversible investigation. Factual uncertainty moved ahead of cosmetic work. The full production record, including the raw excerpt and editorial correction, is in the accompanying worked-example report.
Turn labels into a realistic refresh queue
An audit is useful only when the next action is small enough to schedule. Filter the sheet to approved Now items, then estimate effort. A link replacement may take 15 minutes; a full comparison and consolidation may take several hours.
Use four statuses: queued, researching, editing, and reviewed. Keep publication as a separate editorial state if your CMS uses drafts or scheduled posts. Add a completion note with the sources checked and the date. That history prevents the same uncertainty from being rediscovered next quarter.
Set a review date based on the content, not one universal interval. A timeless worksheet may need little attention. A page with software instructions, regulations, prices, or annual statistics deserves a more deliberate check.
The sheet can also reveal missing connections. If two distinct pages form a supporting pair, add a useful internal link rather than combining them. If you need consistent language while revising several pages, build an AI content style guide before the editing batch.
Where AI audit suggestions go wrong
AI sees only the packet you supply. A title and excerpt can hide the page’s real scope. Thin input produces neat but unreliable labels.
Similarity is another trap. Two pages can share a keyword and still serve different stages of a task. Conversely, different titles may answer the same question. Read the pages before changing URLs or consolidating content.
Analytics can be misread too. Search Console data covers Google Search performance, not every reason a page exists. Direct visits, newsletter links, customer support use, and internal navigation may matter. Recent pages may not have enough history for a fair comparison.
Do not paste a large export into an AI tool without checking what it contains. Query strings, author fields, form exports, or notes may expose information that the audit does not need. Minimize the data first.
Finally, an action label is not implementation. Combining pages can affect links, navigation, and reader expectations. Redirects and removals need technical review, backups, and a clear destination. Keep destructive actions outside automatic workflows.
Human review checklist
Before turning an audit row into work, confirm:
- The URL and title match the live page.
- The page’s reader task is described accurately.
- Dates, statistics, product details, and policy claims were checked at their sources.
- Missing data remains blank rather than guessed.
- Search metrics use the same date range and filters when compared.
- A low-traffic page was not dismissed without considering its purpose.
- Any overlap suggestion is based on reading both full pages.
- Distinct stages of a topic were not mistaken for duplication.
- The proposed action is reversible until a person approves it.
- Redirect, archive, and consolidation work has a technical plan.
- Private or unnecessary data was removed before AI analysis.
- The row has a short reason, owner, priority, and next check.
Questions bloggers ask about content audits
Can I run an AI content audit without Search Console?
Yes. Start with URLs, titles, dates, summaries, link checks, and factual review. Analytics can add context, but they are not required to find outdated claims, unclear reader tasks, or obvious overlap candidates.
How many posts should I audit at once?
For a first pass, 10 to 20 posts is manageable. A small batch lets you fix the column definitions and prompts before applying them to a larger archive.
Should AI decide which posts to delete?
No. Let AI flag possible overlap or weak evidence. A person should read the pages, review their purpose and data, and approve any consolidation, redirect, or removal.
What does “keep” mean in the spreadsheet?
It means no current action is justified by the evidence you reviewed. It does not certify that the page is permanently accurate. Add a future review date when the topic can change.
Audit a small batch before the whole archive
Pick one category and enter ten URLs. Build the fact columns first, run the classification prompt, and correct the labels by hand. Only then add overlap analysis or analytics.
The modest approach is slower than asking AI to judge an entire site in one pass. It is also easier to trust. You will finish with a spreadsheet that shows what you know, what remains uncertain, and which editorial task should happen next.