AI Content Engine · system architecture

How the content factory actually works, layer by layer

A multi-brand SEO content engine. An AI agent does every editorial judgment step — research, planning, writing, reviewing — following written rulebooks. A Node.js engine does everything that must be exact — converting, validating, generating images, publishing to the CMS. Every step reads specific files and writes specific files, so the whole system is an auditable chain of documents. The design principle: separate judgment from determinism, and make rule-compliance a gate that physically blocks bad output.

5brands
6pipeline phases
7master runbooks
34page-type playbooks
9skills (triggers)
13engine commands
3AI agents / article
Claude — skills & agents (the workers) Rulebooks — runbooks, schemas, references (the law) Data files — the artifacts produced & consumed Engine code — deterministic Node.js / scripts External world — web research, image AI, headless CMS

Hover (or tap) any box to light up everything it reads and writes. Solid lines = data flowing between files and workers. Dashed lines = rules, triggers, and enforcement. Arrows point in the direction the information travels.

LAYER 0

Governance — the law every phase obeys

These files are loaded before any work starts. They never produce content themselves; they constrain everyone who does. One of them is not even a document — it is a hook that physically blocks output written without proof the rules were read.

CONTROL TOWERCLAUDE.md · AGENTS.md

The constitution: the exact 6-phase pipeline, the non-negotiable rules, and where every detailed procedure lives. AGENTS.md is its identical twin read by Codex.

loaded every session
DISCIPLINEshared/operating-principles.md

Section S2: no fabrication, confidence labels on every claim, linking rules, punctuation bans, the Adherence Protocol (§S2.19), ignore-word-counts (§S2.20).

ORCHESTRATIONshared/subagent-patterns.md

How Claude splits work across parallel agents: max 5 concurrent, each owns ONE file, Pattern H = the merged write+review pass with one shared research dossier.

TOOL ROUTINGshared/tool-capability-map.md · tool-fallback-reference.md

Which research tool to use for what (SERP, scraping, keyword data) and what to fall back to when one fails or runs out of credits.

METHODOLOGYshared/koray-glossary.md

The semantic-SEO vocabulary (Koray Tugberk Gubur's framework) the whole system is built on: topical maps, entities, contextual coverage.

ENFORCER.claude/hooks/adherence-gate.mjs

A machine gate: it BLOCKS any attempt to save an article or brief until that item's adherence file exists. Skipping the rules is physically impossible, not just discouraged.

machine-enforced
PHASE 1–2

Foundations — learn the brand, then learn every rival

Before a single article is planned, Claude researches the product's own website and the live web to write the brand's "identity card", then does the same for each competitor. Every later phase reads these files in full — they are the single source of who the brand is and what its rivals actually offer.

SKILL/brand-foundation · /product-detail

The triggers. Claude studies the product's live site plus outside research and writes the two canonical brand files.

RUNBOOKrunbooks/brand-foundation-runbook.md

The step-by-step procedure: voice, audience, ideal customer, pains, jobs-to-be-done, competitors, positioning.

SKILL/competitor-foundation

Repeats the same deep research for each competitor. Required before ANY comparison, alternative, review, or "best of" page may exist.

RUNBOOKrunbooks/competitor-foundation-runbook.md

The procedure for profiling a rival: their product, pricing, gaps, and reputation on review platforms.

LIVE WEBSERP · Firecrawl · DataForSEO · G2 / Capterra / Reddit

The outside world. Every fact in the system is retrieved fresh from here in-session — nothing is written from memory.

DATA00-brand/{brand}-brand-foundation.md

THE canonical brand context. Never condensed, never copied — read in full by every downstream phase.

read by phases 3–6
DATA00-brand/{brand}-product-details.md

Feature-by-feature product depth. The authoritative source for any claim about the product itself.

DATA00-brand/{brand}-review-methodology.md

A weighted scoring rubric. Every review, comparison, and "best of" page computes its verdict score from this — no score without it. Also published once as the public "How We Review" page.

gate for evaluative pages
DATA00-brand/competitors/{c}/…-brand-foundation.md + …-product-details.md

One pair of files per competitor. Comparison, alternative, and "best of" pages quote these, never guesses.

PHASE 3

Topical map — the master plan of every page the site will ever have

One spreadsheet row per future page: its exact URL, target search query, page type (out of 39 types), which pages it must link to, and its lifecycle status. It is generated by a Python script — never typed by hand — so the plan is reproducible and self-validating.

SKILL/topical-map

Claude reads the brand + ALL competitor foundations, validates every planned query against the real search results, and edits the builder script.

RUNBOOKrunbooks/topical-map-generation-runbook.md

How to design the map: topic clusters, query validation, internal-link architecture.

SCHEMAschemas/topical-map-schema.md · page-type-inventory.md

The column contract for the CSV, plus the catalog of 39 page types (PT1–PT39) — review, comparison, how-to, what-is, FAQ hub…

BUILDER01-topical-map/_scratch/{brand}-build_map.py

Re-runnable Python script — the map's durable source. To change the map you edit this and re-run it; it validates and writes the CSV atomically, preserving lifecycle columns.

never hand-type the CSV
CSV01-topical-map/{brand}-topical-map.csv

The heart of the system. Every page's proposed_url, query, page type, planned internal links — plus the lifecycle columns every phase updates.

hub file — everything reads it
CSV01-topical-map/{brand}-entity-inventory.csv

Every named thing (laws, tools, concepts) the site must cover, and where. Keeps hundreds of articles consistent about facts and terminology.

rebuilt by review
DOCS…-topical-map-overview.md · …-source-log.md

The human-readable summary of the map's strategy, and the audit trail of every source consulted while building it.

SKILL/csv-to-jsonl

Trigger for the engine's convert step.

ENGINEengine convert (csv-convert.mjs)

Deterministically converts the CSV into the machine-readable spec file the engine's commands consume.

JSONLcontent-plan/{brand}-specs.jsonl

Auto-generated, never hand-edited. One JSON object per page; read by validate, publish, and thumbnails.

PHASE 4

Content briefs — one research-backed blueprint per page

For each map row, Claude runs fresh search-results research and writes a brief: the page's angle, required sections, required entities, and its internal-link plan. In a batch, every brief gets its own independent agent doing its own research — briefs never copy each other.

SKILL/content-brief

The trigger. Reads the map row, brand + competitor context, and the live SERP; writes the brief; then marks the map row briefed.

RUNBOOKrunbooks/content-brief-generation-runbook.md

The universal brief procedure: research phases, validation gates, output format.

17 REFERENCESreference-briefs/PT13…PT33-*.md

Per-page-type playbooks layered ON TOP of the master runbook (both always load): a review brief is structured differently from a how-to brief.

DATA02-briefs/briefs/{id}-brief.md

One blueprint per page: intent, angle, heading skeleton, required entities, link plan, competitor gaps to beat.

PHASE 5+6

Production — research once, write, then adversarially review (Pattern H)

Writing and reviewing run as ONE pass per article with three separate specialist agents. A research agent builds a shared evidence dossier. A writer drafts from that dossier only. A reviewer — a stronger model with a deliberately cold, adversarial handoff — re-judges everything and edits the draft in place. No article ships without surviving this.

SKILL/content-write · /content-review

The triggers. /content-write runs the full merged pass; /content-review alone re-reviews an older article from scratch.

ORCHESTRATORClaude (Opus) — batch conductor

Dispatches at most 5 agents at once, then runs the checks no single-file agent can see: anchor-text variety, link resolution, sibling overlap. Updates the map's lifecycle columns.

RUNBOOKrunbooks/content-writing-runbook.md

The writing law: answer-first sections, sentence/paragraph caps, lists over prose, first-person evidence, Grammarly grammar.

RUNBOOKrunbooks/content-review-runbook.md

Self-contained review law: embeds the SEO, AEO, GEO, and grammar rules verbatim. Bar = 10× information gain, "not longer, better".

17 REFERENCESreference-writers/PT13…PT33-*.md

Per-page-type writing playbooks. The writer follows them; the reviewer audits against the SAME file.

AGENT · SONNETResearch subagent

Fetches the live SERP, AI Overview, full competitor pages, and review-platform evidence ONCE per article into a shared dossier — so writer and reviewer never duplicate the research.

DATA03-content/_scratch/{id}/dossier/

The raw evidence locker: session-fresh captures, not interpretations. The reviewer reuses the captures but re-derives every conclusion itself.

GATE FILE03-content/_scratch/{id}/adherence.md

Proof-of-work: a load manifest (which rulebooks were read) + a conformance attestation (every checklist item ticked). The hook demands it before any output can be saved.

AGENT · SONNETWrite subagent

Drafts the article from the dossier, the brief, and the brand files. Ignores the brief's word counts entirely — length follows content need.

AGENT · OPUSReview subagent

Adversarial editor: re-reads the raw evidence cold, re-judges every fact, beats every competitor page, applies the scoring rubric, and edits the article file in place. Runs the validator last.

DATA03-content/articles/{id}-content.md

The article itself — one file per page, edited in place forever (never forked). Frontmatter carries its URL, entities, and review score.

the product
ENGINEengine validate (validate.mjs)

Deterministic final check: no em/en-dashes, sentence & paragraph caps, lead-in before lists, resolvable internal links, entity minimums. Runs only as review's last step.

CSV02-briefs/{brand}-linking-anchor-inventory.csv

Every internal link's anchor text across the whole site — so hundreds of articles don't all link with the same words. Written ONLY by review; read by briefs for context.

review-owned
ASSETS

Images — Claude plans, Codex paints, the engine finishes

Two visual pipelines share one division of labor: Claude decides WHAT each image should be (never forcing one where a table already works), OpenAI's Codex generates the pixels, and Node.js code applies the brand finish. The three run as separate sequential processes driven by one PowerShell script.

SKILL/article-image-plan

Claude reads a finished article and plans its in-body images: which sections earn one, the type, SEO filename, alt text, and the generation prompt. Skips sections a table or list already serves.

RUNBOOKrunbooks/article-image-planning-runbook.md

The planning law: never force images, never re-analyze an already-planned article.

DATA04-assets/article-images/{id}/plan.json + placement doc

The machine-readable image plan per article — prompts, filenames, alt text, and exactly where each image goes.

DRIVERrun-article-images.ps1

PowerShell conductor: runs plan → render → finalize as three separate processes (Codex is never launched from inside Claude).

EXTERNAL AIOpenAI Codex CLI (codex exec)

The image generator. Receives each plan's prompt non-interactively and renders the raw image.

DATA04-assets/{brand}-brand-guideline.md + brand-assets/

The visual identity extracted from the brand's design file: its primary color, the display typeface, and the logo files used for overlays.

ENGINEengine/src/article-images.mjs

The finisher: picks the right logo variant by measuring the image's luminance, overlays it, and compresses — it never regenerates.

ENGINEengine thumbnails (thumbnails.mjs · codex-runner.mjs)

The blog-card thumbnail pipeline: reads the specs, prompts Codex per page, writes finished thumbnails.

DATA04-assets/thumbnails/

One cover image per article, pushed to the CMS by sync-thumbnails.

PUBLISH

Publishing & reconcile — ship to the CMS, then let truth flow back

The engine converts each reviewed article to the CMS's format and pushes it over the CMS API (idempotent — re-running never duplicates). Then the loop closes: the live CMS and sitemap are the source of truth, and a reconcile command rebuilds the map's status columns from them, so the plan can never drift from reality.

ENGINEengine publish (publish.mjs)

Reads a reviewed article + its spec row and orchestrates the push.

ENGINEmd-to-editorjs.mjs

Converts Markdown into EditorJS blocks + HTML — the two body formats the CMS requires.

ENGINEcms-client.mjs · sync-thumbnails.mjs

The API client: authenticates, creates or updates each post by slug, attaches the thumbnail as the item's image.

CMS APIHeadless CMS — Blog collection

Where articles live once published. Authoritative for item IDs, publish state, and thumbnails.

LIVE SITE{site}/blog/sitemap.xml

The only legitimate source of published_url and publish_date — never guessed, never defaulted to today.

ENGINEengine reconcile (reconcile.mjs)

Rebuilds the map's lifecycle columns from the CMS + local files. Idempotent, only upgrades, never demotes. Run at the start of every session that touches lifecycle.

the anti-drift loop

The lifecycle: five states, one direction

Every row in the topical map is in exactly one state. The state only moves forward, and each transition is owned by one phase. Nothing is published without passing review.

planned→ briefed→ written→ reviewed→ published

In the merged pass, written is only a transient blip — articles normally land straight at reviewed because review runs in the same session. And critically, these status columns are a derived cache, not hand-kept truth: if they are ever lost or clobbered, reconcile rebuilds them from the live CMS and the files on disk.

Who writes what, who reads what

FileWritten byRead by
{name}-brand-foundation.mdPhase 1 (brand-foundation skill), onceEvery later phase, always in full — never summarized into a copy
competitors/… foundationsPhase 2, once per competitorTopical map, briefs, and review for all comparison/review/best pages
{name}-review-methodology.mdOnce per brandReviewer (computes every verdict score); published once as "How We Review"
build_map.py → topical-map.csvThe Python builder (structure); the orchestrator (lifecycle columns only)Briefs, production, engine convert/publish/thumbnails — the hub of everything
entity-inventory.csvBuilder initially; rebuilt only by review from finished articlesBrief and write phases, for cross-article consistency
linking-anchor-inventory.csvReview onlyBriefs (context for the link plan)
{id}-brief.mdPhase 4, one independent agent per briefResearch + write agents (intent and required entities — never as a research substitute)
_scratch/{id}/dossier/Research agent, once per articleWrite agent and review agent (raw captures only; each re-derives its own judgment)
_scratch/{id}/adherence.mdThe producing agent, before its outputThe adherence-gate hook — no gate file, no saved article
{id}-content.mdWrite agent; then edited in place by review, foreverImage planner, validator, publish, reconcile
{name}-specs.jsonlengine convert, from the CSV — never hand-editedEngine validate, publish, thumbnails
plan.json + placement docarticle-image-plan skillPowerShell driver → Codex (render) → article-images.mjs (finish)
Headless CMS + live sitemapengine publish / sync-thumbnailsengine reconcile, which writes truth back into the CSV

The rules that make the quality repeatable

Two rulebooks always load together. Every brief and article obeys its master runbook AND its page-type reference (e.g. PT14-product-review.md). The reference overrides only the structure it names; the master stays authoritative on research integrity and validation. Loading one without the other is a hard error.
Adherence is machine-enforced. Before an agent may save output, it must write a manifest proving which rulebooks it loaded, plus a checklist attestation. A hook physically blocks the file write until that proof exists. Compliance is not on the honor system.
No fabrication, ever. Every claim traces to a source fetched in the current session and carries a confidence label. Publish dates come from the live sitemap; product facts from the product-details file; competitor facts from their foundation files or live pages.
Independence between articles, sharing within one. In a batch, each article's agent does its own ground-up research — no sibling is ever a template. But within one article, research happens exactly once (the dossier) and is shared by writer and reviewer.
Review is adversarial by design. The reviewer is a stronger model with a cold handoff: it reuses the raw evidence but re-derives every judgment, must beat every live competitor page, and edits the same file in place. High-stakes evaluative pages get an extra 2-vote refute pass.
Style is validated deterministically. A code validator — not opinion — enforces the mechanical rules: no em/en-dashes in articles, sentence and paragraph caps, a lead sentence before every list or table, every internal link resolving to a real planned URL, each URL used once per page.
The engine never invents; Claude never publishes by hand. Everything requiring judgment (research, writing, reviewing, image planning) is Claude following rulebooks. Everything requiring exactness (CSV→JSONL, validation, format conversion, API pushes, reconciliation) is deterministic code.

I build AI-powered SEO content systems for early-stage B2B SaaS: full-scale SEO, semantic topical maps, content briefs, content writing, and end-to-end content automation. This map is a sanitized view of the architecture; client names and proprietary methodology are abstracted.

LinkedIn Book a 30-min call salehin.riad96@gmail.com See the dependency map