Every SaaS company with a content team has published an AI content case study. They follow the same template: a headline claiming dramatic speed gains, a vague reference to “quality checks,” and a chart showing traffic that started climbing right around the time they switched tools. What they almost never publish is the operation itself: which site, which articles, what the editor actually changed, and numbers you can audit.
This one is different in the most literal way possible: the experiment is the blog you are reading. Between June 18 and July 14, 2026, we published 25 articles on contentosapp.com — every one drafted by the same 7-agent Contentosapp Studio pipeline we sell, and every one reviewed by one human editor (me) before going live. You can check every claim against the live site: the articles, their dates, their sources. This is the day-25 report. We will update it with ranking data at day 90 — not before, because that data does not exist yet.
- Experiment scope: 25 AI-drafted articles in 25 consecutive days on this very blog — four content clusters plus a comparison batch, all produced with the pipeline’s BYOK mode and reviewed by one human editor.
- Pipeline time: about 9.5 minutes per article on average from brief to finished draft, measured by the pipeline’s own production logs.
- Human editing: 20 minutes per article on average — and what the editor fixed most was not grammar. It was verifying competitor facts against sources and restyling tables.
- Cost per article: about $0.20 in AI usage via our own API key. The plugin is free — the biggest real cost is editorial attention, not tokens.
- Results at day 25: 22 of 25 articles indexed, with first impressions registering in Search Console. Small numbers, stated as small numbers — the ranking story gets told in the day-90 update.
The 25-Article Experiment: Setup and Workflow
The blog is this one — contentosapp.com, a new domain writing about AI content and SEO for WordPress publishers. That detail matters: this is a competitive niche full of established players, not a low-competition test bed. Every brief was structured the same way: a primary keyword, an editorial angle, a reader profile, the article type (pillar or satellite), internal links to sibling posts, and external reference URLs with verified facts in their descriptions. Production ran on Contentosapp Studio in BYOK mode, publishing remotely to this site as drafts — nothing went live without a human pass.
The 7-agent system works sequentially. The Discoverer analyzes what already ranks and what’s missing. The Strategist turns the brief into an SEO outline. The Researcher gathers and validates real sources. The Writer drafts in the configured brand voice. The Editorial Reviewer audits quality, structure, and factual consistency, flagging anything a human should verify. The Visual Designer prepares image prompts and visual structure. The Social Media agent turns the finished article into distribution copy.
| Metric (day 25) | Value | Source |
|---|---|---|
| Articles published | 25 in 25 days (Jun 18 – Jul 14, 2026) | This site’s public archive |
| Avg. pipeline time per article | 9m 28s (10-production sample; range 7m 48s – 13m 07s) | Studio production logs |
| Avg. article length | ~2,400 words | Studio production logs |
| Avg. AI cost per article (BYOK) | ~$0.20 (with image generation; text-only runs less) | Provider billing console |
| Avg. human editing time | ~20 min | Editor’s log (honest estimate) |
| Articles indexed at day 25 | 22 of 25 (the three newest, published this week, still pending) | Google Search Console |
That human-editing line is the number most AI content case studies omit entirely. The delta matters. Total cycle time, not AI generation speed, is the number your business should be measuring.

The Real Numbers: Time, Cost, and What the Editor Changed
Competing case studies report AI generation time the way car ads report horsepower. It sounds impressive and tells you almost nothing. So here is the honest cost structure. AI usage in BYOK mode — our own API key across all seven agents — averaged about $0.20 per article: the plugin is free, and BYOK mode has no per-article fee or markup. The single biggest token cost is image generation — text-only articles run meaningfully cheaper.
The real cost is editorial attention. The human pass averaged 20 minutes per article. Price that at whatever your time is worth — the point is that it does not disappear, and any case study that only reports AI generation time is hiding the biggest line item. If you want to model these numbers against word-cap and seat-fee tools, our AI content cost per article breakdown maps it in detail.
What the Editor Actually Changed (and What Surprised Us)
Across 25 articles, the editing pass settled into a predictable pattern — and it was not fixing grammar. The pipeline’s prose consistently arrived publish-ready at the sentence level. The real work fell into four repeating categories.
First: restyling tables. The pipeline generates standard tables; we replace them with our house-styled comparison tables on every article that has one.
Second: verifying competitor facts against sources. This is the highest-stakes category. In one comparison article, the draft stated a competitor plan limit that appears nowhere on that vendor’s pricing page — a plausible-looking number pulled from training data instead of the referenced source. We caught it by searching the vendor’s page for the literal number. The correction made the article’s argument stronger, because the verified numbers were more favorable than the invented ones.
Third: cutting weak citations. When the pipeline’s web research surfaces a low-authority source (a random listicle, an unrelated case study), it will sometimes anchor a claim to it. Every external link gets checked against the brief’s reference list; anything that did not come from the brief gets scrutiny.
Fourth: small editorial curation — adding a cross-link the brief asked for, softening an unsourced generalization, syncing the FAQ schema with edited answers.
Results at Day 25 — and What We Will Report at Day 90
This is where most case studies would show you a hockey-stick chart. We can’t — the blog is 25 days old, and pretending otherwise would defeat the purpose of this piece.
What we can report at day 25: 22 of 25 articles are indexed — the three newest, published this week, are still in the queue. Search Console shows just over 2,000 impressions and exactly 3 clicks site-wide so far.
Those are small numbers, stated as small numbers. A new domain does not outrank entrenched competitors in its first month, with or without AI. What the first 25 days actually validate is the operating model: a solo founder shipped 25 researched, sourced, internally-linked articles in 25 days, at a marginal software cost of about $0.20 per article, without the quality collapsing into slop.
The ranking story gets told at day 90, in an update to this post: which articles reached page one, at which keyword difficulties, and how the lightly-edited articles performed against the heavily-edited ones. Bookmark it — the update ships in October 2026.

Frequently Asked Questions
Does AI-generated content actually rank on Google in 2026?
Ask us at day 90 — seriously. This blog is 25 days old, and honest ranking data takes a quarter, not a month. What we can say at day 25: 22 of 25 articles are indexed, none show any sign of spam demotion, and Google’s own guidance targets low-quality content, not AI-assisted production. The 90-day update to this post will publish the ranking table.
How long does it realistically take to produce one WordPress article using an AI pipeline, including editing?
We’ll publish audited averages in the day-90 update, but the day-25 numbers are: about 9.5 minutes of pipeline time on average (measured by the production logs) plus roughly 20 minutes of human editing. The editing time varied widely — comparison articles with competitor pricing to verify took the longest; workflow guides often needed only a light pass.
What does a human editor actually need to fix in AI-drafted content before publishing?
In our experience across these 25 articles: restyling tables to the house standard, verifying every competitor number against the vendor’s own page, cutting citations to weak sources that automated research occasionally surfaces, and syncing the FAQ schema after edits. Grammar and structure almost never needed work — the drafts arrive publish-ready at the sentence level. The one thing the pipeline cannot supply is lived experience; that layer comes from the editor.
Is AI content on WordPress penalized by Google’s helpful content system?
None of the 25 articles in this experiment received a visible demotion attributable to AI generation. Google has repeatedly stated that its systems target low-quality, unhelpful content — not AI content specifically. The practical risk is publishing AI output that fails the “who wrote this and why should I trust them” test. That failure looks the same whether a human or a machine produced it. The editorial layer in this pipeline — particularly the fact-verification pass — is the mechanism that addresses this risk directly.
How much does it cost to produce one AI-assisted blog post end-to-end?
The software side: the plugin is free, and in BYOK mode the AI usage billed by our own provider averaged about $0.20 per article across this batch. The real cost is editorial time — about 20 minutes per article. Price that at your own hourly rate: the point of this case study is that the editing line item is real and belongs in any honest calculation.
What keyword difficulty range is realistic for AI content to rank within 90 days?
That is exactly what the day-90 update of this post will answer, with a table: which of the 25 articles reached page one, at which keyword difficulties. At day 25 the honest answer is “too early to say” — indexing is underway and impressions are starting to register, but ranking claims this early would be projections, and this piece exists specifically to avoid those.
Running the same 25-article experiment on your own blog will not produce identical numbers. The niche, domain authority, brief quality, and editor experience all move the outputs. But this gives you a template with auditable line items: a public archive of 25 articles you can inspect, production times measured by the pipeline itself, and an editing pattern documented category by category — with the ranking table to follow in the day-90 update. That is a more actionable starting point than any case study claiming generalized “efficiency gains.” If you want to run the same workflow, start with 10 articles in a single topical cluster, log every edit the human makes, and check rankings at 30, 60, and 90 days. The pipeline’s results become credible once you can audit your own version of them.
Transparency note: every article on this blog — including this case study — is drafted by the same 7-agent Contentosapp Studio pipeline described here, and edited by a human before it goes live.

