Copywriting prose creator

01What is it?
Provides prose engineer. What sets it apart is how it narrows content production into one specific workflow rather than a broad, generic prompt.
02Inputs
Context for content production: your goals, audience, constraints, and any source material the skill asks for.
03Output
A ready-to-use result for content production: the analysis, copy, or recommendations the agent produces.
Install-only

Install as a package

Installs this one skill package for your coding agent, including any supporting files that skill ships with — not every skill in the repository. Read the tutorial.

Terminal
$ npx skills add samber/cc-skills --skill copywriting-prose-creator

Skill instructions

The instruction file for this skill. The skill also includes other files you need to install to use it.

SKILL.md

Persona: You are a prose engineer. Prose is reproducible craft, not art — codify lexicon, syntax, rhythm, structure, and voice markers so any writer (human, ghostwriter, or AI) can hit the same fingerprint.

Thinking mode: Use ultrathink for every BUILD and ADAPT invocation. Prose codification synthesizes multi-input artifacts (SOUL.md + TONE.md + corpus + interview), arbitrates conformity-vs-differentiation against category defaults, and projects rules onto multiple supports. Shallow reasoning produces generic guides that flatten into LLM-default register — the exact failure mode this skill exists to prevent.

Modes:

  • BUILD — fresh PROSE.md from SOUL.md + TONE.md + discovery interview (sequential)
  • ADAPT — port an existing PROSE.md to a new channel grouping (sequential)
  • AUDIT — corpus analysis to surface current prose patterns before codification (parallel sub-agents when corpus > 50 pieces)

Copywriting Prose

Produces PROSE.md: a brand-specific prose guide that codifies how a brand writes, independent of what it feels like. Prose is the observable craft a forensic linguist could measure on a page — sentence length, clause depth, lexicon, parallelism, signature moves. Tone is the emotional posture, handled separately. Two brands with identical tones can have non-interchangeable prose; that is what this guide captures.

The slogan: tone is the music, prose is the score. This skill codifies the score.

Inputs and outputs

ArtifactRoleProducer
SOUL.md (optional)Storyteller archetype, mission, POVsibling skill
TONE.md (optional)Emotional posture (NN/g 4 dimensions)samber/cc-skills@copywriting-tone-of-voice-creator
Existing PROSE.mdSource for ADAPT modethis skill
Content corpusSource for AUDIT modebrand's CMS / blog / social archives
PROSE.mdOutputthis skill

DESIGN.md (visual identity) sits in the same register but is out of scope. PROSE.md becomes the system-prompt substrate for downstream writers: samber/cc-skills@linkedin-ghostwriting, samber/cc-skills@substack-ghostwriting, samber/cc-skills@technical-article-writer, samber/cc-skills@press-release-writer.

Channel groupings

Per project convention, channels are treated as four generic groupings, not as platform-specific surfaces. Platform-specific quirks (LinkedIn's algorithm, Substack's paywall) live in the writer skills, not in PROSE.md.

GroupingCovers
Long-form articlesBlog posts, pillar pages, evergreen essays, technical deep-dives, opinion essays (Substack, Medium, dev.to, own blog — same group)
Social postsLinkedIn, X, Bluesky, Threads, TikTok captions, Mastodon
Email & newsletterNewsletter issues, transactional, drip sequences, lifecycle emails
Marketing copyLanding pages, ad copy, press releases, podcast show notes, video scripts, sales decks

BUILD workflow

Phase 0 — Detect inputs

Look in the working directory (and common locations like ./brand/, ./content/, ./docs/) for SOUL.md, TONE.md, prior PROSE.md, and any content corpus. If SOUL.md or TONE.md is missing, surface this — these artifacts feed directly into Phases 1 and 3, and proceeding without them forces inline assumptions that lock the prose guide to a sketch instead of the brand's actual archetype.

If missing, offer two paths:

  1. Invoke the sibling skill first (samber/cc-skills@copywriting-tone-of-voice-creator for TONE.md). Why: TONE.md captures the brand's emotional posture across the four NN/g dimensions; without it, prose rules drift into tone territory and become unfalsifiable.
  2. Capture archetype and tone minimally inline (Phase 1 interview adds a short addendum). Pragmatic for one-off prose audits.

If a content corpus exists, offer to run AUDIT mode first — empirical patterns beat invented ones every time.

Phase 1 — Discovery interview

Use AskUserQuestion in 2–3 batches. Skip any field already supplied by SOUL.md, TONE.md, or prior conversation context. Wait for answers before proceeding — assumptions in the interview compound into a wrong prose guide that downstream writers will faithfully reproduce.

Required fields (full battery in references/discovery-questions.md):

  • Brand mission (one sentence)
  • Category posture: conformist, adjacent, challenger, outsider
  • Audience: reading age, expertise (Layperson / Practitioner / Expert), locale, language(s), patience
  • Author archetype (read from SOUL.md if present, else ask): journalist · engineer · founder · NGO advocate · politician · consultant · executive · community lead · artist · researcher
  • Objective per channel: awareness · engagement · lead · signup · retention · advocacy
  • Distribution channels: long-form · social · email · marketing copy (multiSelect)
  • Constraints: legal, regulatory, brand safety, confidentiality
  • Cultural context: HQ locale vs audience locale, language(s) of operation
  • Tone of voice (if TONE.md missing): NN/g four dimensions quick-pick — funny↔serious · formal↔casual · respectful↔irreverent · enthusiastic↔matter-of-fact

Phase 2 — Category detection and deep-research routing

Match the brand to one of the 11 covered categories. Load the playbook from references/category-playbooks.md — it carries category-specific defaults for mean sentence length, lexicon, signature structures, anti-patterns, and reference brands.

#Category
1B2B (SaaS / enterprise tech)
2B2C (consumer products)
3Consumer brand (lifestyle / DTC)
4Non-corporate / NGO / non-profit
5Consulting / professional services
6Product-led (makers, indie hackers, dev tools)
7Industry (manufacturing, deep-tech, industrial)
8Volunteering / community / association
9Personal branding (per-principal)
10Politics / advocacy / public figures
11Internal corporate communication

Uncovered context → delegate research. When the brand sits clearly outside the 11 categories — for example religion / faith-based, defense / military, healthcare / pharma regulated, finance regulated, legal practice, cultural institutions (museum / opera / theater), educational institutions, government communications, intelligence services PR, esports, adult content, crypto / web3, niche luxury, fashion / beauty editorial, kids / edutainment, agritech, climate / environmental advocacy with policy posture — surface the gap and invoke samber/cc-skills@deep-research to map the category's prose conventions before codifying. Why: category playbooks compress 30+ pieces of corpus evidence per category; codifying without that substrate produces guides that read like generic LLM output.

For personal branding the same logic applies per principal: a corpus capture of 60–90 minutes of the principal's recorded speech plus prior writing is required before codifying. Generic personal-branding rules produce ghostwritten posts that read like every LinkedIn founder.

Phase 3 — Codify the five layers

Codify each layer in order. Each rule needs a why — bare prescriptions without rationale fail the moment a writer hits an edge case. Detail rules and examples in references/five-layers.md.

  1. Lexicon — use/avoid A–Z (50–200 entries), terminology table, jargon ladder per channel, acronym policy, naming conventions, foreign-word policy, technical depth scale (Layperson / Practitioner / Expert)
  2. Syntax — mean sentence length target (category default, ±2), distribution targets (≤10% of sentences ≥25 words; ≥15% ≤8 words for rhythm), clause depth, active voice default with exception list, parallelism rules, paragraph length and architecture
  3. Rhythm — cadence variance target (σ ≥ 6 words per 100-word window), breath points (one ≤8-word sentence every 3–5 sentences), repetition policy, callbacks, list patterns, white-space cadence
  4. Structure — opening hook types (cross-ref samber/cc-skills@copywriting-hooks), closing types (cross-ref samber/cc-skills@copywriting-cta), transitions, headings (sentence case, frontloaded), subheadings, lists, asides, quotations, citations, blockquotes, reader positioning (Gardner's far↔close psychic distance: default per channel, shift-signal words, when to close for conversion)
  5. Voice markers — 5–12 signature moves, signoffs, recurring metaphors, idioms, taboos, intentional tics (all rationed; unrationed markers collapse into self-parody)

Diagnose the corpus before locking the targets:

  1. wc -w and a sentence-length distribution script (see references/audit-tools.md) — establish current mean and σ before declaring targets
  2. Hemingway readability against a sample of 5 pieces — sanity-check the reading age claim from Phase 1
  3. grep -i for each candidate banned word in the existing corpus — confirm the brand actually drifts toward it before banning

Phase 4 — Punctuation and formatting policies

Two non-negotiable tables.

Punctuation policy — declare a position on each: em dash, en dash, semicolon, colon, ellipsis, parentheses, italics, bold, single/double quotes, exclamation marks, brackets, hyphens (compound modifiers), Oxford comma, capitalization (sentence vs title case). Defaults and rationing tables live in references/five-layers.md (references/five-layers.md#punctuation).

Formatting policy — heading hierarchy (H1 once, H2 sections, H3 sub-sections, max H4 in technical docs only), bullet rules (3–7 items, parallel grammar, leading sentence), numbered lists (only when order matters), code blocks (language tag, line cap), images (caption + alt text), callouts (rationed), tables (only for 2D relationships), links (frontloaded link text — never "click here", "learn more", "read more"). Why frontloaded link text: scannability and accessibility; screen readers extract link lists out of context.

Phase 5 — Channel-specific overrides

For each in-scope channel grouping (see table above), produce a CHANNEL section in PROSE.md with deltas on sentence length, paragraph length, hook types, closing types, formatting, and CTA pattern. Pull the transformation rules from references/channel-adaptation.md.

Generic groupings keep PROSE.md portable: when a brand adds a new platform within a grouping (e.g. moves from Threads to Bluesky), the overrides hold without re-codification.

Phase 6 — Cultural and linguistic adaptation

  • English variant: declare US / UK / international English (spelling, punctuation, date format)
  • French ↔ English: list the few French words permitted in English text (raison d'être, savoir-faire) and forbid others without translation; conversely declare English loan-words accepted in French (le marketing, le briefing) vs taboo
  • False cognates: éventuellement ≠ eventually, actuellement ≠ actually, important often ≠ important; full list in references/multilingual.md
  • Transfer budgets: cut 20% of words FR→EN, pad 20% EN→FR — French rewards longer sentences, English brand prose favors shorter
  • Locale conventions per channel grouping: French LinkedIn cadence differs from US conventions in formality, paragraph length, first-person use
  • Accessibility and inclusion: bias-free language section (people-first, singular "they", preferred pronouns)

For multilingual brands: one PROSE.md per language, not a translated single guide. Maintain a mapping document of shared pillars and divergent rules.

Phase 7 — Anti-LLM countermeasures

The dominant prose-drift risk in content factories is convergence on LLM-default register. Codify rules LLMs do not follow by default — that is the durable defense.

Full inventory in references/anti-patterns.md. Headline patterns:

  • Lexical tells: delve, leverage, crucial, robust, underscore, navigate (as transitive metaphor), seamlessly, vibrant, dynamic, embark, foster, harness
  • Structural tells: tricolons in series ("X, Y, and Z"), summative closers ("In conclusion…"), colon-titles ("The Future of X: A New Paradigm"), bullet-list overuse, hedged claims without source
  • Punctuation tells: em-dash overuse (single signal — not proof; see Ann Handley's published rebuttal); ellipsis outside quotation
  • Formula constructions: "It's not just X, it's Y" · "Picture this:" · "Imagine a world where" · "What if I told you" · "Whether you're a seasoned X or a curious newcomer" · "In the realm of" · "Navigating the landscape of"

Diagnose LLM drift quantitatively:

  1. grep -c -iE 'delve|leverage|crucial|robust|underscore' across the corpus — frequency ≥1 per 500 words is a strong tell
  2. Sentence-length σ < 4 across a 100-sentence window — uniformity is a stronger tell than any single lexical signal
  3. n-gram comparison between the brand's pre-AI corpus and post-AI corpus — divergence in top trigrams flags drift

Detection is unreliable as a single source of truth. Use these as triage, not verdict. The Stanford HAI / Liang et al. (2023) work showed GPT detectors misclassify TOEFL essays by non-native English writers at headline rates above 60%. Treat any single signal as suspicion, not proof.

Phase 8 — Render PROSE.md

Use the hybrid template in references/prose-md-template.md:

  1. Narrative sections for each layer + policy (the why and the how)
  2. Do/don't tables as an annex (the quick-reference scan layer)
  3. Sample bank: ≥10 before/after pairs, ≥3 exemplar pieces if provided, hook bank and closing bank cross-referenced from samber/cc-skills@copywriting-hooks / @copywriting-cta
  4. Cross-references to TONE.md and SOUL.md (read together, not in isolation)
  5. Versioning footer: semver, date, owner, changelog stub

ADAPT workflow

Take an existing PROSE.md and project it onto a new channel grouping.

  1. Read the existing PROSE.md.
  2. Ask the user: target channel grouping (long-form / social / email / marketing copy), and optionally a specific platform within the grouping for tighter overrides.
  3. Compute the transformation delta from references/channel-adaptation.md: sentence-length cut or grow factor, paragraph break frequency, hook style adjustment, CTA fit, formatting overrides.
  4. Emit a CHANNEL OVERRIDE — <grouping> section appended to PROSE.md, or a standalone PROSE-<grouping>.md if the user prefers a separate artifact. Why offer both: content teams that publish across many channels prefer one master file; ghostwriting agencies handling a single channel prefer per-channel files.
  5. Cross-reference back to the original PROSE.md for fields unchanged.

AUDIT workflow

Extract current prose patterns from a corpus before codifying. Empirical patterns beat invented ones.

  1. Take the corpus (folder of .md / .txt or list of URLs).
  2. For corpora > 50 pieces, parallelize: spin up to 5 sub-agents via the Agent tool, splitting the corpus by date range, channel, or author. Each agent reports back with the same metrics. Why parallel: sequential reading on a 200-piece corpus is slow and runs out of context; parallel sub-agents read independently and synthesize.
  3. Compute (per references/audit-tools.md):
    • Mean sentence length and distribution
    • Top 50 lexemes, top bigrams and trigrams
    • Banned-word and AI-tell frequency
    • Em-dash count per 1,000 words
    • Opening pattern map (first 50 words of 30 pieces, side by side)
    • Closing pattern map
  4. Run an adversarial reading pass on 3–5 representative pieces — challenge the assumption that they work. Mark every sentence that doesn't earn its place, every unanswered reader question, every moment authority collapses, every paragraph where a reader would disengage. See references/audit-tools.md (references/audit-tools.md#adversarial-reading) for the methodology.
  5. Sort findings into four buckets: signature (recurring, distinctive, working) · default (recurring, generic, neutral) · noise (inconsistent, accidental, weak) · liability (recurring, actively harming credibility or engagement — the adversarial pass surfaces these).
  6. Produce AUDIT-MEMO.md (5–10 pages: quantitative tables + qualitative annotated samples + "keep, kill, differentiate" summary). Feed into BUILD Phase 3.

Output format

PROSE.md
├── Cover (brand, version, owner, last updated, status)
├── Purpose (200 words: who it is for, how to use, what it does not cover)
├── Prose Pillars (one page, 5–8 falsifiable pillars)
├── Voice vs. Tone note (one paragraph)
├── 1. Lexicon (narrative + do/don't annex)
├── 2. Syntax
├── 3. Rhythm
├── 4. Structure
├── 5. Voice Markers
├── 6. Punctuation Policy
├── 7. Formatting Policy
├── 8. Channel Overrides (one section per in-scope grouping)
├── 9. Cultural & Linguistic Adaptation
├── 10. Anti-LLM Countermeasures
├── 11. Sample Bank (before/after, exemplars, anti-exemplars, hook bank, closing bank)
├── 12. Ghostwriting Addendum (per principal — optional)
├── Annex A: Do/Don't quick reference (all layers, scannable)
└── Changelog

A complete PROSE.md is 20–60 pages depending on category coverage and channel scope. Resist the urge to maximize length — Siemens reduced their brand guidelines from 2,750 to 250 pages because enforceable density beats exhaustiveness. Aim for the density that an editor can apply line by line; cut anything an editor cannot turn into a concrete edit.


Reference files (load on demand)

FileWhen to read
discovery-questions.md (references/discovery-questions.md)During Phase 1 interview
five-layers.md (references/five-layers.md)During Phase 3 codification
category-playbooks.md (references/category-playbooks.md)During Phase 2 after category detection
channel-adaptation.md (references/channel-adaptation.md)During Phase 5 and all ADAPT invocations
anti-patterns.md (references/anti-patterns.md)During Phase 7 and AUDIT mode
multilingual.md (references/multilingual.md)During Phase 6 when brand operates in EN/FR
prose-md-template.md (references/prose-md-template.md)During Phase 8 render
brand-atlas.md (references/brand-atlas.md)During Phase 2 archetype matching
audit-tools.md (references/audit-tools.md)During AUDIT mode and Phase 3 corpus diagnosis

Disclaimer

This skill is not exhaustive. The 11 category playbooks compress a much larger landscape — refer to the brand's own corpus, the linked frameworks (Mailchimp, IBM Carbon, GOV.UK, Microsoft, Atlassian, Buffer), and canonical references (Ann Handley Everybody Writes, Joseph Williams Style, Roy Peter Clark Writing Tools, Margot Bloomstein Trustworthy) when the playbook does not cover the situation. For uncovered categories, invoke samber/cc-skills@deep-research and feed its output back into BUILD Phase 2. Prose guides decay; a PROSE.md not re-audited every 12 months is a snapshot, not a living document.

If you encounter a bug or unexpected behavior, open an issue at https://github.com/samber/cc-skills/issues.


Supporting file: references/anti-patterns.md

Anti-Patterns and AI Tells

The dominant prose-drift risk in content factories is convergence on LLM-default register. This file inventories the patterns to ban explicitly in PROSE.md and to check for in AUDIT mode.

Refresh yearly. Lexical and structural tells evolve as models change.

Lexical tells

Words and phrases that appear at elevated frequency in LLM-generated text. Roger J. Kreuz (The Conversation, 2025) documented "a dramatic increase in relatively uncommon words, such as 'delves' or 'crucial,' in articles published in scientific journals over the past couple of years."

Single-word tells

WordCommon in LLM output asReplacement strategy
delve"let's delve into…""examine", or just say the thing
leverage"leverage X to achieve Y""use", or restructure to active verb
crucial"this is crucial"drop or be specific about why
robust"robust framework / system"name the property concretely
underscore"this underscores the importance of…""shows" / "proves" / "means"
navigate"navigate the complexities of…"drop the navigation metaphor; say the thing
seamlessly"integrates seamlessly"replace with what it actually does
vibrant"vibrant community"give a concrete signal of vibrancy
dynamic"dynamic landscape"be specific about what changes
embark"embark on a journey"drop the journey metaphor
foster"foster collaboration""build" / "support" / "encourage"
harness"harness the power of"drop the harnessing metaphor
myriad"a myriad of options""many" / specific number
tapestry"rich tapestry of"drop the tapestry metaphor
paradigm"new paradigm""model" / "approach" — usually drop
pivotal"pivotal moment"be specific about why it pivoted what
holistic"holistic approach"name the components
synergizen/abanned in any register
utilize"utilize""use"
facilitate"facilitate""help" / "let" / "enable"
commence"commence operations""start" / "begin"
furthermoresentence openerreplace with "and" / drop / restructure
moreoversentence openerreplace with "also" / drop
notwithstanding"notwithstanding the…""despite"

Phrase tells

  • "It's not just X, it's Y" — the formula construction; LLMs emit this 5–10× human baseline rate
  • "In conclusion,…" — summative closer; humans usually skip it
  • "It's important to note that…" — hedging that adds nothing
  • "It's worth noting that…" — same
  • "Crucially,", "Notably,", "Importantly,", "Essentially," — sentence-opener hedges
  • "Picture this:" / "Imagine a world where" / "What if I told you" — manufactured-curiosity openings
  • "Whether you're a seasoned X or a curious newcomer" — audience-segmenting filler
  • "In the realm of" / "Navigating the landscape of" — empty scene-setting
  • "Unlock the power of" / "Dive into" / "Buckle up" / "Let's dive in" — hype openers
  • "From the comfort of your own home" — copywriting cliché
  • "At the end of the day" — empty connector
  • "Last but not least" — banned in lists
  • "In today's fast-paced world" / "In an era of" — universally banned opener
  • "Hope this helps!" — chatbot signature

French phrase tells

  • "Dans un monde en constante évolution" — universally banned French opener
  • "Plongez dans…" — equivalent of "dive into"
  • "Découvrez comment…" — over-used in French AI output
  • "Par ailleurs,…" / "Notamment,…" — over-used sentence openers
  • "Il est crucial de…" — French equivalent of "it is crucial to"
  • "À l'heure du tout-numérique" / "À l'ère de l'IA" — universally banned openers
  • "N'hésitez pas à…" — chatbot-flavored closing

Structural tells

Isabel Al-Dhahir (Verdict, 2024): "AI models also overuse tricolons, a rhetorical device that consists of a series of three parallel words, phrases, or clauses."

PatternWhy it's a tellCounter
Tricolons in series ("X, Y, and Z")LLMs default to threes everywhere; humans vary list lengthsLimit lists of 3 to 1 per 1000 words; use 2-item or 4-item lists elsewhere
Colon-titles ("The Future of X: A New Paradigm")Over-represented in LLM headlinesUse single-clause titles; drop the colon
Summative closers ("In conclusion,…")Humans skip this in 80%+ of piecesEnd on the thing itself; no recap
Bullet-list overuseLLMs convert prose to bullets reflexivelyAudit: bullet-density per piece. Cap.
Hedged claims without source"It is widely believed that…" / "Many experts agree…"Force citations or remove
Symmetric paragraph lengthsUniform 4-sentence paragraphs across the pieceForce variance via the rhythm rules
Anaphora overuseRepeated openers in clustersReserve anaphora for closings only
"Not only X, but also Y"LLM-favored coordinationUse "X. And Y." or restructure

Punctuation tells

The em-dash debate is contested. Ann Handley's published rebuttal (The Em Dash Is NOT an AI Tell, annhandley.com) argues it is unreliable as a single signal. Treat as suspicion, not proof.

Punctuation patternStrength of signal
Em-dash density > 1 per 200 wordsWeak — humans vary widely
Ellipsis outside direct quotationModerate
Parenthetical asides > 1 per 200 wordsWeak
Em dash + tricolon + summative closer in same pieceStrong

Statistical tells

GPTZero (Edward Tian) frames detection as perplexity (how predictable to a language model) plus burstiness (sentence-length variance).

  • Low perplexity correlates with AI generation. Hard to measure without tooling.
  • Low burstiness = uniform sentence length. Measurable: σ < 4 words per 100-sentence window suggests automation. σ ≥ 6 is the human baseline target.

Detection unreliability

Liang et al. (2023), GPT detectors are biased against non-native English writers (Patterns, Cell Press; Stanford HAI): "These platforms incorrectly labeled more than half of the essays as AI-generated, with one detector flagging nearly 98%" of TOEFL essays by non-native English writers.

Operational consequence: use detectors as triage, not verdict. Codify rules LLMs cannot follow by default — that is the durable defense, not a detector arms race.

Counter-measures for content factories

  1. Codify rules that LLMs cannot follow by default: idiosyncratic lexicon, named taboos, sentence-length variance targets, specific voice markers.
  2. Use the PROSE.md as the AI system prompt; supplement with the brand's published corpus as few-shot examples.
  3. Editor reviews 100% of AI-assisted content with the audit scorecard.
  4. Run regular n-gram comparison between the brand's pre-AI corpus and post-AI corpus to detect lexical drift.
  5. Invoke samber/cc-skills@humaniseur-fr (for French) or an equivalent humanizer skill as a final pass on AI-assisted drafts. Note: humanizers do not replace the prose guide — they scrub the patterns the guide already banned.

Cliché opener inventory

A list to embed in the brand's taboos section. These all immediately disqualify the writer:

English:

  • "In today's fast-paced world…"
  • "Have you ever wondered…?"
  • "Did you know…?"
  • "What if I told you…?"
  • Dictionary opener played straight ("Productivity, defined as…")
  • "In this article, I'll discuss…"
  • "I'm not an expert, but…"
  • Three rhetorical questions in a row
  • "Imagine waking up…" without a specific scene
  • "Hot take:", "Unpopular opinion:"
  • "At [Company], we believe…"
  • "Recently,…" without a specific date
  • "You're not alone."
  • "We've all been there."
  • "Buckle up,"
  • Misattributed Einstein / Seneca / Confucius / Bouddha quotes

French:

  • "À l'heure du tout-numérique"
  • "À l'ère de l'IA"
  • "Dans un monde où…"
  • "Vous êtes-vous déjà demandé…?"
  • "Dans cet article, nous allons voir…"
  • "Je ne suis pas spécialiste mais…"
  • "Voici la vérité que personne ne veut entendre…"
  • "Récemment,…" sans date précise
  • "Chez [Entreprise], nous pensons…"
  • "Cher lecteur," (in newsletter)
  • "Les études montrent que…" sans source
  • "90% des gens…" sans source

Diagnose

When auditing a piece for AI tells:

  1. grep -c -iE 'delve|leverage|crucial|robust|underscore|seamlessly|navigate|harness|foster|embark|myriad|tapestry|holistic|paradigm|utilize|facilitate|commence' — count single-word tells. > 1 per 500 words = strong tell.
  2. Count em dashes per 1000 words. > 5 = check the context.
  3. Compute sentence-length σ on a 100-sentence sample. σ < 4 = robotic cadence.
  4. Search for the formula constructions ("It's not just X, it's Y"; "Whether you're a seasoned X"). Any hit = rewrite.
  5. Check tricolon density per 500 words. > 3 = AI-favored.

Supporting file: references/audit-tools.md

Audit Tools

For AUDIT mode and for Phase 3 corpus diagnosis. The goal: compute objective prose metrics on a corpus so that the prose guide's targets are empirical, not invented.

These tools do not produce prose recommendations directly. They surface signals; the model interprets them and codifies the rules.

Readability

Hemingway Editor

Browser tool at hemingwayapp.com (or the desktop / pasteable web version). Highlights:

  • Sentences hard to read (yellow) and very hard (red)
  • Passive voice
  • Adverbs
  • Complex phrases ("utilize" → "use")
  • Reading grade level

Usage: paste 1000 words. Read off the grade level. Targets per category:

CategoryTarget grade
B2C / consumer brand6–9
B2B SaaS9–12
NGO / nonprofit7–10
Industry / deep-tech12–16
Consulting11–14

Why grade level matters: it correlates with sentence length, syllable count, and Latinate ratio. A grade level mismatched with audience expertise predicts drop-off.

Plain Language Commission

For UK English specifically. Maps to the same metrics with slightly different targets.

Banned-word enforcement

Vale

Vale (vale.sh) is the dominant prose linter. It applies YAML rules to Markdown / plain text and flags violations.

Sample Vale rule (banned word):

extends: existence
message: "Banned word: '%s'. Use a plain alternative."
level: error
tokens:
  - delve
  - leverage
  - crucial
  - robust
  - underscore
  - seamlessly
  - navigate
  - harness
  - foster
  - embark
  - myriad
  - tapestry
  - paradigm
  - utilize
  - facilitate
  - commence

Place under .vale/styles/Brand/BannedWords.yml. The brand can extend with category-specific bans (e.g., "synergize" for consulting).

Why Vale, not LanguageTool, for banned words: Vale is built for prose-rule enforcement; LanguageTool is built for grammar correction. Different tools, different jobs.

LanguageTool

LanguageTool (languagetool.org or the open-source self-host) catches grammar errors, typos, and weak constructions. Strong on:

  • Subject-verb agreement
  • Passive voice (alerts; does not auto-rewrite)
  • Wordiness suggestions
  • Repeated words

Configuration: import the brand's banned-word list as a custom dictionary; whitelist intentional brand terms (the lowercase Innocent, the all-caps Apple SHORTCUTS, etc.).

Grammarly Business

Same job class as LanguageTool, commercial. Stronger UI for distributed writers. Custom dictionary supports the brand's banned and required terms. Browser plugin enforces in-flow.

Quantitative metrics

Mean sentence length and distribution

A short Python snippet for a quick audit:

import sys
import re
from statistics import mean, stdev

text = sys.stdin.read()
# Split on sentence terminators, ignoring abbreviations
sentences = re.split(r'(?<=[.!?])\s+', text)
sentences = [s.strip() for s in sentences if s.strip()]
lengths = [len(s.split()) for s in sentences]

print(f"Sentences: {len(sentences)}")
print(f"Mean length: {mean(lengths):.1f} words")
print(f"Std dev: {stdev(lengths):.1f}")
print(f"Min / max: {min(lengths)} / {max(lengths)}")
print(f"≥ 25 words: {sum(1 for l in lengths if l >= 25)} ({100*sum(1 for l in lengths if l >= 25)/len(lengths):.1f}%)")
print(f"≤ 8 words: {sum(1 for l in lengths if l <= 8)} ({100*sum(1 for l in lengths if l <= 8)/len(lengths):.1f}%)")

Run on each piece in the corpus. Aggregate per channel.

Type-token ratio

Vocabulary diversity. Lower TTR = more repetition (often a sign of LLM-flat lexicon). Compute as unique_words / total_words on a 1000-word window. Healthy human B2B prose lands at 0.40–0.55; LLM-default sits at 0.30–0.42 due to favored vocabulary clustering.

n-gram frequency

For lexical drift detection. Compare two corpora (e.g., pre-AI vs post-AI):

from collections import Counter

def ngrams(text, n=3):
    tokens = text.lower().split()
    return Counter(' '.join(tokens[i:i+n]) for i in range(len(tokens) - n + 1))

pre = ngrams(open('pre_ai_corpus.txt').read())
post = ngrams(open('post_ai_corpus.txt').read())

# Trigrams that surged
diff = [(t, post[t] - pre[t]) for t in post]
diff.sort(key=lambda x: -x[1])
for t, d in diff[:30]:
    print(f"{d:+5d}  {t}")

Top surging trigrams flag the brand's drift vectors.

Banned-word frequency

# Count banned-word hits across a corpus
grep -roEic 'delve|leverage|crucial|robust|underscore|seamlessly|navigate|harness|foster|embark|myriad|tapestry|holistic|paradigm|utilize|facilitate|commence' ./corpus/ | sort -t: -k2 -nr

Hits per 500 words is the relevant rate.

Reading-aloud audit

The lowest-tech tool, often the most revealing. Read the piece aloud:

  1. Mark every breath point.
  2. Mark every sentence where you stumble or have to re-read.
  3. Mark every monotonous stretch (more than 4 consecutive medium sentences).
  4. Mark every accidental rhyme or alliteration.

A trained editor catches in 5 minutes what statistical tools miss. Codify the read-aloud audit as part of the per-piece QA checklist for high-stakes content (pillars, executive op-eds, keynotes).

n-gram comparison for ghostwriting

For ghostwriting voice-matching audits: compute n-gram overlap between the principal's authentic writing/speech corpus and the ghostwriter's drafts. Low overlap on signature n-grams = the ghostwriter has flattened the principal's idiolect.

This is the operational test for the "thinking translation problem" — when the ghostwriter reproduces the client's sentence structure and vocabulary but misses the actual operational insight only the client could have.

Web-based readability tests

For a quick sanity check without local tooling:

  • Hemingway Editor (hemingwayapp.com) — grade level, complex sentences, passive voice, adverbs
  • Datayze Sentence Length Checker (datayze.com/sentence-length-checker) — distribution histogram
  • WebFX Readability Test (webfx.com/tools/read-able) — multiple readability scores

Vale + CI integration

For content factories at scale, wire Vale into the editorial CI:

# .vale.ini
StylesPath = .vale/styles
MinAlertLevel = warning

[*.md]
BasedOnStyles = Brand, Microsoft, write-good

Run on every pull request to the content repo. Fail the build on Vale errors. Why: the prose guide is enforced by tooling, not by editor fatigue. Editors should review judgment calls, not catch banned words.

Adversarial reading

Quantitative tools surface signals. Adversarial reading surfaces what the numbers miss: the sentences that don't earn their place, the moments authority collapses, the reader questions that go unanswered.

Core posture: the writer already believes the draft works. Challenge that assumption. Read to find what fails, not to confirm what succeeds.

Protocol (per piece)

Read the piece once without stopping. Then re-read and mark:

  1. Dead weight — sentences or phrases that could be deleted without information loss. Count them. A ratio above 15% signals a draft, not a final piece.
  2. Authority collapse — claims that invite "says who?", statistics without sources, analogies that don't hold, jargon used to signal expertise rather than convey meaning.
  3. Reader dropout points — paragraphs where a reader would disengage: slow accumulation with no payoff, five consecutive medium-length sentences with no breath point, transitions that require re-reading.
  4. Unanswered questions — the "so what?" question every factual claim generates. If the paragraph raises a question and the next paragraph doesn't resolve it, the structure is broken.
  5. Distance incoherence — sudden shifts in psychic distance (see five-layers.md § 4.11 (five-layers.md)) with no structural reason (e.g., a close-second-person hook that snaps to third-person-corporate in paragraph two).

Critique dimensions for brand prose

Adapted from fiction critique methodology (haowjy/creative-writing-skills@prose-critique):

DimensionFor brand proseKey question
StructurePiece-level coherence, pacing, payoff mechanicsDoes each section earn its length? Does the opening pay off at the close?
VoicePOV stability, dialogue effectiveness, implicit meaningDoes the brand voice hold across the full piece, or drift mid-article?
ProseSentence-level craft, rhythm, word repetition, descriptor-to-narrative ratioAre there more than 3 consecutive sentences of the same type?
Brand personaMotivational consistency of the brand characterDoes the brand's stated archetype match the brand's actual prose behavior?
ContinuityFactual accuracy, claim consistency within the pieceDoes the piece contradict itself across sections?

Sorting findings

Map each finding to a bucket:

  • Signature — recurring and working; codify as rules to preserve
  • Default — recurring and neutral; decide whether to keep or differentiate
  • Noise — accidental and inconsistent; no action needed
  • Liability — recurring and harmful to credibility or engagement; codify as explicit prohibitions

The liability bucket is what adversarial reading surfaces that metrics miss.

Limits of automation

These tools surface signals; none replaces editorial judgment. A piece can pass every quantitative check and still read flat — because the voice markers are missing, or the structure is wrong, or the hook doesn't earn the close. Use audit tools as the first filter, then ship to human editorial review for the rest.


Supporting file: references/brand-atlas.md

Brand Atlas

A short, opinionated catalogue of brands whose prose is identifiable in a blind test. Use during Phase 2 archetype matching: ask "which of these does the brand most resemble?" and "where does it want to differentiate?"

Each entry: one-line summary of the diagnostic feature.

Consumer brands

Mailchimp — short, declarative, sentence-case, contractions, Oxford comma, plain Anglo-Saxon verbs, sparing humor, never patronizing. Diagnostic: sounds like a colleague writing an email at 3pm.

Innocent Drinks — ultra-short sentences (frequently 4–8 words), product-as-narrator first person, lowercase brand name, hand-written-feeling asides, fruit puns rationed. Diagnostic: a sentence that ends in "Win." or a parenthetical aside in product voice.

Patagonia — long, journalistic, magazine-feature openings; named human protagonists (Chilean fishermen, female carpenters, Japanese sake brewers); data points embedded in narrative; declarative ethical claims. Diagnostic: a story about a named person with a place name in the dek.

Liquid Death — maximalist horror-genre register, mock-tabloid microcopy, all-caps headlines, sentence fragments as exclamations, satirical "About" pages. Diagnostic: a product description that reads like a B-movie pitch.

Oatly — handwritten-style asides, faux-naïve voice, self-deprecation, parenthetical meta-commentary on its own marketing.

Apple — terminal sentence economy and sentence-case discipline.

B2B / SaaS brands

Stripe — footnoted, narrative-memo prose. Patrick Collison structures emails like research papers with footnotes. Diagnostic: a paragraph that pre-empts an objection in parentheses or a footnote. Stripe defaults to writing over slide decks across the entire company.

Linear — terse, declarative, no marketing adjectives, single-sentence value propositions, present-tense product claims, dense product screenshots in lieu of explanation. Diagnostic: a sentence with no adjective.

Slack — friendly imperatives, microcopy as small talk, sentence-case headings, contractions, frequent second person. Diagnostic: a UI string that begins with a verb and feels like a coworker.

Notion — didactic-but-warm, lists everywhere, second-person, "you can…" constructions, definitions before features. Diagnostic: help-doc cadence even in marketing copy.

Basecamp / 37signals — opinionated, contrarian, short paragraphs, frequent one-sentence paragraphs, dialectical structure (claim/counterclaim/resolution). Diagnostic: a one-sentence paragraph that contains an opinion.

Mailerlite — clear, didactic, list-heavy, second-person, low jargon, screenshot-rich. Diagnostic: a numbered list that is the article.

Social media voices

Wendy's social — punchline-first, sentence fragments, callbacks to previous posts, no hashtags. Diagnostic: a reply that is funnier than the prompt.

Duolingo social — chaos register, lowercase, ironic threats, owl-as-narrator. Diagnostic: a post that would be a fireable offense at any other brand.

Reference frameworks

The published style guides worth reading in full before codifying:

Nielsen Norman Group, The Four Dimensions of Tone of Voice (Moran, 2016) — funny↔serious, formal↔casual, respectful↔irreverent, enthusiastic↔matter-of-fact. Use as sanity check; not a prose rule.

Mailchimp Content Style Guide (styleguide.mailchimp.com) — most-copied open template. Take the structure (Voice and tone, Grammar and mechanics, Writing about people, Writing for accessibility).

GOV.UK Style Guide — the discipline of plain language. Targets reading age 9 across the site; maintains a strict A–Z banned-word list. Take the discipline of a maintained A–Z list.

Microsoft Writing Style Guide (learn.microsoft.com/style-guide) — sentence case in headings, Oxford comma, contractions allowed, "Write short, simple sentences." Mean target 15–20 words per sentence.

IBM Carbon Design System (carbondesignsystem.com/guidelines/content/overview) — codifies voice as a nine-attribute checklist. Take the rare technique of writing voice as a checklist an editor can apply line by line.

Atlassian Design System (atlassian.design/content) — "be bold, be optimistic, be practical, with a wink" tetrad, paired with the internal "Voicify" sliding-scale tool. Take the situational-flex model for newsletters and product-led content.

Buffer Style Guide (buffer.com/resources/style-guide) — voice attributes "relatable, approachable, genuine, inclusive". Product copy rules unusually prescriptive: "Invite the customer to take an action. Never command." Take the prescriptive product-copy block.

Canonical writing books

For the brand owner and editor to read once, not the writers:

  • Ann Handley, Everybody Writes (2nd ed., 2022) — the 17-step "Writing GPS", the "ugly first draft" doctrine, "Create reading momentum."
  • William Zinsser, On Writing Well — the pruning doctrine ("clutter is the disease of American writing"). Operationalize as a 20% cut rule on every draft.
  • Strunk & White, The Elements of Style — rule 17 ("Omit needless words"), rule 11 ("Use the active voice"). Overrideable defaults.
  • Joseph Williams, Style: Lessons in Clarity and Grace — the character-action principle (subjects should be characters, verbs should be actions). The most useful syntactic rule for technical writers escaping nominalizations.
  • Roy Peter Clark, Writing Tools — Tool 2 ("Order words for emphasis": strongest word at the end, second strongest at the start) and the parallelism toolkit.
  • Heath brothers, Made to Stick — the SUCCES heuristic for openers and closers, especially Concreteness.
  • Margot Bloomstein, Content Strategy at Work and Trustworthy — the BrandSort message-architecture method, the prerequisite to a prose guide.
  • Nicole Fenton and Kate Kiefer Lee, Nicely Said — the simple style-guide template at the end of the book is the minimum viable prose guide.

Corpus linguistics methods

For quantitative audits, the relevant methods:

  • Type-token ratio — vocabulary diversity per piece
  • Mean sentence length distribution — central rhythm signal
  • n-gram frequency — top bigrams and trigrams; drift detection
  • POS-tag profiles — verb/noun/adjective ratios; detects nominalization drift
  • Banned-word frequency — direct enforcement metric

These are the only way to detect prose drift in a large content operation. See audit-tools.md for the actual tools.

How to use this atlas in Phase 2

  1. Ask the brand owner which 1–3 brands they admire most (not necessarily in their category).
  2. Ask which 1–3 brands they actively want to avoid sounding like.
  3. Map both sets to the entries above where possible; for unrecognized brands, run an inline corpus skim.
  4. Codify the "differentiate" axis: which of the loved brand's signature moves to borrow; which of the avoided brand's defaults to ban.

The atlas is not exhaustive. A brand that does not resemble any entry here is interesting — codify what they actually do, then file it under a new entry for next time.


Supporting file: references/category-playbooks.md

Category Playbooks

Eleven playbooks. Each one compresses corpus evidence for the category — load the relevant playbook in Phase 2, apply its defaults in Phase 3.

Format per playbook: Optimize for · Prose characteristics · Anti-patterns · Lexicon · Syntax · Rhythm · Structure · Reference brands.

For categories not in this file, invoke samber/cc-skills@deep-research to map the category's prose conventions before codifying.

1. B2B (SaaS, enterprise, tech)

  • Optimize for: clarity + credibility
  • Prose characteristics: declarative-default; concrete examples within 100 words of any abstract claim; numbers, dates, customer names; precise terminology; explicit connectors (because, therefore, in contrast); structured arguments
  • Anti-patterns: vague benefit-speak ("transform your business"); adjective stacking ("powerful, scalable, intelligent"); long sentences with passive subjects ("It is believed that…"); thought leadership without an actual thought
  • Lexicon: technical terms permitted at a known specificity (latency, throughput, TCO, MRR) with glosses for non-expert readers; ban marketing intensifiers (best-in-class, world-class, cutting-edge); use product names exactly as registered
  • Syntax: mean 14–18 words; allow longer for technical claims with multiple qualifications; subject-verb proximity (no buried verbs); active default, passive permitted in technical impersonal claims
  • Rhythm: alternate short claim with longer evidence; one breath sentence (≤ 8 words) every 4–6 sentences; H2/H3 every 200–300 words
  • Structure: TL;DR or summary up top; inverted pyramid; numbered lists for procedures; tables for comparisons; explicit headings ("How it works", "Why it matters", "What to do")
  • Reference brands: Stripe (footnoted memos), Linear (terse declaratives), Notion (didactic-warm), Atlassian (clear product prose)
  • France-specific: French B2B SaaS audiences in English expect lower tolerance for hyperbole and higher tolerance for technical depth than US norms

2. B2C (consumer products)

  • Optimize for: emotional clarity + memorability
  • Prose characteristics: shorter sentences; second person; concrete sensory verbs (taste, feel, see); product as protagonist; benefit-first leads; one idea per paragraph
  • Anti-patterns: B2B jargon (solutions, platforms); features without sensory translation; long sentences; passive constructions; abstract nominalizations
  • Lexicon: Anglo-Saxon verbs over Latinate; no acronyms without explanation; product names always capitalized as registered
  • Syntax: mean 10–14 words; sentence fragments permitted in body for emphasis; imperatives common in CTAs
  • Rhythm: punchy openings; one-sentence paragraphs allowed; lists kept to 3 items (Rule of Three)
  • Structure: lead with the feeling, then the feature; testimonials and reviews integrated; product photography as part of the structural rhythm
  • Reference brands: Innocent Drinks, Mailchimp consumer copy, Apple consumer pages

3. Consumer brand (lifestyle, DTC)

  • Optimize for: distinctiveness + memorability
  • Prose characteristics: high voice-marker density; idiosyncratic signature moves; willingness to break grammar rules deliberately (Innocent's lowercase i, Oatly's run-on asides); brand-as-character voice; meta-commentary on its own marketing acceptable
  • Anti-patterns: imitating Innocent without earning it (Nick Asbury's "Wackywriting" critique); cuteness without substance; ironic distance that masks lack of conviction
  • Lexicon: brand-owned coinages encouraged (Liquid Death's "Murder Your Thirst", Oatly's "Wow no cow!"); strong opinions encoded in vocabulary
  • Syntax: deliberately varied; fragments common; one-sentence paragraphs frequent; rule-breaking permitted but rationed
  • Rhythm: highly variable cadence; surprise sentence lengths; punchlines at the end
  • Structure: micro-content first (packaging, social); pillars often manifestos rather than how-to
  • Reference brands: Innocent Drinks, Liquid Death, Oatly, Patagonia (purpose-led variant), Cards Against Humanity (irreverent variant)
  • Caveat: this register collapses without a real product point of view. Liquid Death works because the product (canned water in beer-style cans) is itself a thesis

4. Non-corporate / NGO / non-profit

  • Optimize for: dignity + clarity
  • Prose characteristics: humans named with consent, not anonymous "beneficiaries"; data embedded in narrative, not opposed to it; specific places and dates; first-person testimonials with full attribution where possible; ethical care in describing vulnerable populations
  • Anti-patterns: poverty porn / suffering aesthetics; abstract nouns of suffering ("hunger", "injustice") without grounding; corporate-speak imported from for-profit ("stakeholders", "engagement", "ROI of empathy"); patronizing framings ("giving voice to" instead of platforming)
  • Lexicon: people-first language ("people experiencing homelessness", not "the homeless"); strengths-based vocabulary (Charity: Water — focus on hope, not guilt); avoid Latinate abstractions in favor of Anglo-Saxon verbs; maintain a banned list for stigmatizing terms
  • Syntax: medium sentence length (15–20 words) to accommodate context; active voice default with named agents; passive when protecting privacy
  • Rhythm: narrative cadence; alternating personal scene and aggregate data; explicit "from one person to many" structure
  • Structure: scene → person → context → systemic claim → call to action; named impact ("37 wells in 12 villages, serving 4,800 people"); credit lines for community partners
  • Reference brands: Charity: Water (hope-not-guilt), Médecins Sans Frontières (clinical-but-human), Oxfam (campaigning), Patagonia (purpose-led)
  • France-specific: French NGO register is more abstract and lexically Latinate than English equivalents; English-language adaptations should cut 20% of abstract vocabulary

5. Consulting / professional services

  • Optimize for: authority + accessibility
  • Prose characteristics: structured arguments with named frameworks; data with citations; institutional "we have observed" or partner-level "I have observed"; concrete client situations (anonymized where required); explicit caveats and conditions
  • Anti-patterns: McKinsey-pastiche ("In our experience, leading organizations…") without specifics; consulting bingo (synergies, optimization, transformation); claims without evidence; pseudo-data ("studies show")
  • Lexicon: declared framework terms (capitalize when proper noun, lowercase otherwise); banned consultancy jargon list (synergize, operationalize, leverage as verb); precise hedges ("we observed in 7 of 12 engagements", not "often")
  • Syntax: longer sentences acceptable (18–22 words mean) given expert audience; subordination acceptable; parenthetical conditions common
  • Rhythm: slower cadence than B2B SaaS; long paragraphs of argument (5–8 sentences) acceptable in pillars; broken by data callouts or numbered findings
  • Structure: executive summary mandatory; situation–complication–resolution (Minto Pyramid) common; numbered findings; explicit "implications for X" sections
  • Reference brands: McKinsey Quarterly (institutional), Bain Insights (data-driven), Stripe Press (technical-cultural), BCG Henderson Institute (academic-adjacent)

6. Product-led (makers, indie hackers, dev tools)

  • Optimize for: utility + product as evidence
  • Prose characteristics: present-tense product claims; screenshots and code blocks as part of the prose; "show, don't tell" enforced (Linear's approach); CTAs are product demos, not contact forms; documentation tone leaking into marketing copy (positively)
  • Anti-patterns: enterprise marketing imported into product-led contexts ("revolutionize your workflow"); long preambles before showing the product; benefits without screenshots
  • Lexicon: feature names treated as proper nouns; verbs that map to UI actions (click, select, configure); no industry buzzwords unless category requires
  • Syntax: very short, mean 10–14 words; imperatives in how-to content; declaratives for product claims
  • Rhythm: claim, screenshot, claim, screenshot; very high signal-to-noise
  • Structure: above-the-fold product image plus one-sentence claim; problem-solution-proof below; final CTA = try the product
  • Reference brands: Linear, Notion, Vercel, Plausible, Raycast

7. Industry (manufacturing, B2B industrial, deep-tech)

  • Optimize for: precision + technical credibility
  • Prose characteristics: specifications cited exactly; standards referenced (ISO, IEC, ASTM); named engineers and facilities; metric units; long-form technical depth in pillars with executive-summary openings
  • Anti-patterns: consumer marketing language (amazing, incredible); soft claims without spec (high performance); breathy purpose-marketing pasted onto industrial products
  • Lexicon: domain-specific terminology required, not avoided; standard abbreviations as norms (kW, MPa, IP67); banned consumer intensifiers; full part numbers and model designations
  • Syntax: long sentences tolerated (up to 25-word means in technical sections); passive voice acceptable in scientific impersonal register; precise qualifications ("at 25°C and 1 atm")
  • Rhythm: slower; longer paragraphs; data tables interrupt prose
  • Structure: GE Reports model is the journalistic benchmark — named protagonist, named place, named challenge, measurable outcome; brand mentioned sparingly
  • Reference brands: GE Reports (journalistic storytelling), Siemens (industrial minimalism), Anthropic and OpenAI research blogs (deep-tech variant — long-form, hedged, mission-framed)
  • Deep-tech variant: research-blog prose hedges epistemically ("preliminary results suggest"); cites primary sources; reserves superlatives; embeds equations and diagrams

8. Volunteering / community / association

  • Optimize for: belonging + actionability
  • Prose characteristics: inclusive pronouns (we, our community); concrete next actions per piece (a meetup, a contribution, a vote); volunteer-authored voices given billing; gratitude as a structural element, not closing platitude
  • Anti-patterns: corporate-charity register imported from NGOs; insider jargon excluding newcomers; assumed knowledge ("as you know from last month's meeting"); event posts missing basics (date, place, time, who can come)
  • Lexicon: declared insider terms with one-line glosses for newcomers; banned gatekeeping vocabulary; inclusive terms (newcomers vs "noobs")
  • Syntax: short, conversational; second person plural often appropriate; imperatives for calls to action ("RSVP by Friday")
  • Rhythm: brisk; bullet lists for logistics; narrative for stories from members
  • Structure: who-what-when-where-why in the first 100 words for events; member spotlights as recurring format; "how to get involved" as standard closing
  • Reference brands: open-source community newsletters (Rust Foundation, Kubernetes); volunteer-run conferences (PyCon, GopherCon); local-chapter newsletters

9. Personal branding (individual creators, founders, executives)

  • Optimize for: authentic voice fingerprint + cumulative point of view
  • Prose characteristics: first person used deliberately; signature openings; short paragraphs on social (1–3 sentences); idiolect preserved (specific filler words, recurring connectors, characteristic metaphors); opinions clearly stated
  • Anti-patterns: ghostwritten prose that smooths the principal's idiolect into generic LinkedIn cadence; thought leadership without thoughts; performative vulnerability; AI-template structure (hook + three bullets + question CTA)
  • Lexicon: principal-specific. Maintain the principal's actual word list (favorite verbs, characteristic adjectives, terms they refuse to use)
  • Syntax: principal-specific. Some founders use long sentences (Paul Graham); others use short (Naval Ravikant). Match the principal's mean sentence length within ±2 words
  • Rhythm: principal-specific. Capture 60–90 minutes of recorded speech to extract
  • Structure: post archetypes (lesson + story, contrarian opinion, framework, reaction to industry event); each archetype has an established cadence per principal
  • Diagnostic test: a blind reader should be able to identify the principal from a paragraph
  • Ghostwriting addendum (mandatory for per-principal codification):
    1. Corpus capture: 60–90 min of recorded speech + all prior writing (emails, posts, op-eds)
    2. Idiolect analysis: mean sentence length, top 50 lexemes, characteristic openings, filler phrases (you know, the thing is), preferred connectors, recurring metaphors, frequent stories, banned topics
    3. Codify: 10 signature openings · 5 signature closings · 3 recurring stories with retelling cadence · banned topics · preferred connectors · average post length · line-break convention
    4. Calibration: 3 test posts; principal reviews; iterate until "this sounds like me" with no edits
    5. Quarterly review with principal

10. Politics / advocacy / public figures

  • Optimize for: clarity of position + memorability + ethical durability
  • Prose characteristics: declarative-default; concrete promises and concrete constraints; rhetorical devices (anaphora, tricolons, contrast) used deliberately but rationed (overuse signals AI or amateurism); explicit acknowledgement of opposing views before refutation; consistent message across speeches, posts, op-eds
  • Anti-patterns: hedging that obscures position; insider procedural language; attack-only register; recycled boilerplate from prior campaigns; superlatives ("most important election of our lifetime") that erode credibility through repetition
  • Lexicon: campaign-specific declared vocabulary (signature words and phrases that recur); avoidance of opposition's framing language; inclusive pronouns balanced with specific named groups; banned terms list to prevent gaffes
  • Syntax: shorter than other categories (mean 12–16 words); parallelism enforced in promises and value statements; tricolons reserved for keynote moments
  • Rhythm: oratorical when intended for speech (read-aloud testing mandatory); written rhythm for op-eds; both should match the principal's natural cadence
  • Structure: Aristotle's ethos-pathos-logos still operative. For speeches: hook (named constituent), thesis, three reasons, opposition acknowledged, peroration. For op-eds: argument-led, evidence-rich, concrete proposal
  • Reference: Gettysburg Address (272 words, three paragraphs, ten sentences) as the canonical short modern model; Amnesty International and Human Rights Watch for advocacy NGO benchmarks
  • Ethical guardrail: stick to what can realistically be accomplished. Misleading voters wins applause in the short term but erodes trust in the long run. Invoke samber/cc-skills@deep-research for jurisdiction-specific legal constraints (campaign finance disclosure, defamation thresholds, electoral codes) before codifying
  • Note: per-principal customization required, similar to personal branding addendum

11. Internal corporate communication

  • Optimize for: trust + comprehension + actionability
  • Prose characteristics: plain language; named owners and named deadlines; "what changes for you" sections; explicit acknowledgement of uncertainty during change; consistent voice from leadership across channels (all-hands, intranet, email, Slack)
  • Anti-patterns: corporate euphemism for layoffs / restructuring ("synergies", "right-sizing", "transitioning colleagues out"); buried bad news (the real news in paragraph 6); jargon-as-power ("strategic alignment workstream"); ChatGPT-flavored "I am pleased to announce" boilerplate
  • Lexicon: employee-facing terms (colleagues, teammates) over HR-coded ones (resources, headcount); concrete role names over abstract function names (engineering team over engineering organization); accessibility-first (avoid acronyms unique to one division)
  • Syntax: mean 14–18 words; second person ("you", "your team") in change comms; imperatives for action items
  • Rhythm: bullet-heavy for change comms; narrative for context and rationale; clear delineation between "what" / "why" / "what changes for you" / "what to do next"
  • Structure: lede paragraph carries the news (no throat-clearing); FAQ section for change announcements; named contact for follow-up questions; explicit timeline
  • Reference: Patrick Collison's leaked Stripe internal memos style (research-paper structure with footnotes); Basecamp / 37signals all-hands updates (contrarian-but-warm); Slack's internal comms playbook (transparency by default)
  • Critical: this category collapses into HR-speak more reliably than any other. The prose guide must explicitly ban the corporate-comms cliché set (cascading communication, leveraging synergies, key stakeholders, going forward, at this time)

Multi-category brands

A brand may sit in two categories simultaneously (e.g., a B2B SaaS that markets like a consumer brand — Notion, Linear). In that case, write one PROSE.md per audience segment, not a blended guide. A blended guide collapses to the lowest common denominator and loses both audiences. Maintain a mapping document for shared pillars and divergent rules.

Disclaimer

These playbooks compress the dominant patterns of each category as observed in the source research. They are starting points, not endpoints. Audit the brand's actual corpus (via AUDIT mode) before locking the category defaults — empirical patterns beat compressed ones.


Supporting file: references/channel-adaptation.md

Channel Adaptation

Transformation rules between the four generic channel groupings. Used in BUILD Phase 5 (to populate channel-overrides sections in PROSE.md) and in ADAPT mode (to project an existing PROSE.md onto a new channel).

Generic groupings keep PROSE.md portable: adding a new platform within a grouping does not require re-codification. Platform-specific quirks (LinkedIn's algorithm, Substack's paywall, X's reply economy) live in the downstream writer skills, not here.

The four groupings

GroupingCovers
Long-form articlesBlog posts, pillar pages, evergreen essays, technical deep-dives, opinion essays (Substack web post, Medium, dev.to, own blog — same group)
Social postsLinkedIn, X / Twitter, Bluesky, Threads, TikTok captions, Mastodon
Email & newsletterNewsletter issues (Substack email, ConvertKit, Beehiiv), transactional, drip sequences, lifecycle
Marketing copyLanding pages, ad copy, press releases, podcast show notes, video scripts, sales decks

Channel deltas

Each row is a transformation rule. When ADAPT mode projects a PROSE.md from one grouping onto another, walk this table top-down.

DimensionLong-formSocialEmail & newsletterMarketing copy
Mean sentence length14–18 words (B2C 10–14; consulting 18–22)8–12 words10–14 words8–12 words
Paragraph length2–5 sentences, max 80 words1–3 sentences, frequent breaks1–2 sentences1–3 sentences
Hook styleScene · contrarian · stat · concrete detail · questionBold claim · direct problem · concrete detail · counterintuitivePersonal frame · curiosity gap · open loopPromise · direct problem · authority · benefit
Body structureTES, PEEL, inverted pyramidHook → re-hook → ABT (And/But/Therefore) → CTAPersonal lede → 3–5 paragraphs → P.S.Above-fold claim → proof → CTA
List useBulleted lists within prose, max 7 itemsNumbered or bulleted, heavy use, max 5 itemsSparse — break with bold insteadBullet-heavy for features
Heading frequencyEvery 200–350 words (H2/H3)None (line breaks instead)Light — bolded section headersOne H1, optional H2 per section
ClosingPractical next step · callback · restated stakeSpecific reply prompt · no open-ended questionsSingle primary CTA in P.S. or final paragraphDirect action button + risk reversal
Voice marker densityModerate (1–2 signature moves per piece)High (1 marker per post — the brand's "tell")Moderate (personal opener as marker)Low (consistency over personality)
Reading time3–15 min15–60 sec2–5 min30–90 sec
Tone registerVoice unchanged, tone matter-of-fact to enthusiasticVoice unchanged, tone casual-confidentVoice unchanged, tone personal-intimateVoice unchanged, tone direct-confident

Transformation rules (long-form → other channels)

The most common ADAPT direction. Pillar articles are usually the source of truth; other channels derive.

Long-form → social post (200–300 words; one platform-specific tweet/post)

  1. Extract the single most counterintuitive claim from the article.
  2. Lead with it as the hook (re-engineer per samber/cc-skills@copywriting-hooks).
  3. Cut sentence length to ~60% of the long-form mean.
  4. Break paragraphs to 1–3 sentences.
  5. Add line breaks for scannability.
  6. CTA: specific reply prompt, comment, link to long-form, or no-CTA. Avoid open-ended questions — they reduce action.
  7. Strip qualifiers ("often", "in general", "typically") — social rewards confidence.

Long-form → newsletter issue (500–1500 words)

  1. Open with the personal or editorial frame ("This week I've been thinking about…", "A reader emailed me…").
  2. Compress argument to 3–5 paragraphs, one idea per paragraph.
  3. Include one named data point within the first 200 words.
  4. Single primary CTA — in P.S. or final paragraph. Newsletter readers reward restraint.
  5. Cut images sparingly — many email clients block them by default. Lead with strong cover image if used.
  6. Subject line: declarative, specific. Avoid clickbait — newsletters live on trust.

Long-form → marketing copy (landing page, ad, press release)

  1. Strip narrative scaffolding (anecdotes, scene-setting). Marketing copy lives above the fold.
  2. Lead with the promise or direct problem. The reader must know within 3 seconds whether to keep reading.
  3. Translate features → benefits → outcomes.
  4. Add risk reversal (guarantees, social proof, testimonials).
  5. Single CTA per page. Multiple CTAs split conversion.
  6. Banned in marketing copy unless brand-specific: rhetorical questions, hedges, scene openings.

Long-form → email & newsletter drip sequence

  1. Split the article's arguments across N emails — one argument per email.
  2. Each email opens with a hook that pulls forward from the previous one.
  3. Closing of each email teases the next (curiosity gap, open loop).
  4. Final email contains the primary CTA. Earlier emails build trust without asking.
  5. Maintain consistent sender voice across the sequence — drip sequences live on continuity.

Cross-channel transformation rules (other directions)

Social → long-form (post-as-seed)

When a social post performs well, expand it. Rules:

  1. The post becomes the lede of the long-form piece.
  2. Add 3–5 sections that defend or extend the original claim.
  3. Add evidence the original post could not carry (data, named cases, citations).
  4. Replace the social CTA with a long-form closing (callback or restated stake).

Email → social (newsletter-as-thread)

  1. Pick the strongest single argument from the newsletter.
  2. Recast as a thread or single post.
  3. Preserve the personal frame — newsletter readers expect intimacy; social readers reward authenticity.

Marketing copy → social

Generally avoid. Marketing copy that reads like marketing copy on social produces low engagement. If forced, strip all sales language and focus on a single insight from the marketing piece.

Hooks per channel

Cross-reference samber/cc-skills@copywriting-hooks for the full hook catalog. Per channel grouping, the permitted hook types are:

GroupingPermitted hooks
Long-formScene · contrarian · curiosity gap · concrete detail · stat · question · historical analogy · time anchor · authority
SocialBold claim · direct problem · concrete detail · contrarian · stat · conditional · pattern interrupt
Email & newsletterPersonal confession · curiosity gap · open loop · conditional · personal frame
Marketing copyPromise · direct problem · authority · stat · benefit

CTAs per channel

Cross-reference samber/cc-skills@copywriting-cta for the full CTA archetype catalog. Per channel grouping:

GroupingCTA archetypes
Long-formPractical next step · callback · restated stake · transitional asset (lead magnet)
SocialSpecific reply prompt · link to long-form · profile visit nudge
Email & newsletterSingle P.S. CTA · reply prompt · paid-tier tease (if applicable)
Marketing copyDirect action button · book a call · free trial · pricing page

Anti-patterns per channel

ChannelAnti-patterns
Long-formSlow scene-setting in technical pieces; closing with generic "What do you think?"; lists-as-article (when prose would do)
SocialHashtag spam; emoji as personality substitute; threading what should be one post; open-ended questions as CTA
Email & newsletterSalesy subject lines; multiple competing CTAs; ignoring preview text; long paragraphs (email clients render them as walls)
Marketing copy"Click here", "Learn more"; feature lists without translation; testimonials without faces and names; multiple CTAs per page

When ADAPT mode emits a separate file

Two output options for ADAPT mode:

  1. Inline channel override section appended to PROSE.md (preferred when channels share most rules)
  2. Standalone PROSE-<grouping>.md (preferred when a channel has substantial divergence, e.g., a B2B SaaS with a wildly different consumer brand for one product line)

Ask the user which they prefer in ADAPT Phase 2.


Supporting file: references/discovery-questions.md

Discovery Questions

Full intake battery for BUILD Phase 1. Use AskUserQuestion in 2–3 batches; skip any field already supplied by SOUL.md, TONE.md, or prior conversation. Treat as a kickoff session checklist — a 90-minute interview fills the critical fields; subsequent sessions backfill the rest.

Organized by domain. Mandatory fields are marked [M]; the rest are nice-to-have and improve the guide but do not block it.

1. Brand / entity

  • [M] Mission in one sentence (founder's words and marketing's words, side by side)
  • [M] Message architecture: 9–12 prioritized attributes (Bloomstein BrandSort or equivalent)
  • [M] Brand voice owner operationally (CMO, founder, head of content)
  • [M] Brand-to-category posture: conformist · adjacent · challenger · outsider
  • Brand age, and whether prose still matches current stage

2. Audience

  • [M] Reading age and literacy level (GOV.UK aims at age 9; B2B SaaS commonly at 14–16; expert pubs at 18+)
  • [M] Domain expertise: Layperson · Practitioner · Expert
  • [M] Cultural context: US · UK · France · EU · global · other
  • [M] Language(s) the audience reads in, and fluency
  • Professional reading habits (newsletters, publications they trust)
  • Patience level: consumer scrolling vs professional researching

3. Existing content

  • Top 10 highest-performing pieces of the last 12 months, by channel
  • Bottom 10 (failure modes are diagnostic)
  • Where voice is consistent, where it drifts
  • Which pieces were ghostwritten, agency-produced, or AI-assisted
  • What the support / sales team says about how customers describe the brand voice

4. Competitors and category conventions

  • 3–5 competitors whose prose is studied (or copied) internally
  • Default category register (e.g., enterprise B2B "thought leadership")
  • Conventions to conform to vs break, and the cost of breaking each

5. Distribution

  • [M] Channels in scope (multiSelect: long-form articles · social posts · email & newsletter · marketing copy)
  • Cadence per channel (pieces/month)
  • Channel-specific constraints (character limits, SEO requirements, deliverability)
  • Typical length per channel

6. Writers and operations

  • Who writes, in what mix (employees, freelancers, agencies, ghostwritten principals)
  • How briefing is done (templates, briefs, voice notes)
  • Review workflow (number of rounds, who has veto)
  • Editorial calendar planning horizon
  • Tools (Google Docs, Notion, CMS, AI assistants)

7. Content goals

  • [M] Per channel, primary KPI: awareness · engagement · lead · signup · retention · advocacy
  • How prose changes if KPI is awareness vs conversion vs retention
  • Relationship between prose distinctiveness and conversion (sometimes inverse for compliance-heavy categories)

8. Constraints

  • Legal: regulated claims, disclaimer requirements, IP/trademark conventions
  • Regulatory: industry-specific (finance, health, defense, pharma)
  • Compliance: GDPR consent language, accessibility (WCAG 2.2 AA)
  • Brand safety: topics that are off-limits
  • Confidentiality: what cannot be discussed publicly

9. Cultural context

  • [M] Locale of brand HQ vs locale of audience
  • [M] Language(s) of operation
  • Cultural taboos and sensitivities (regional, religious, political)
  • Geopolitical positioning (does the brand take positions on global events?)

10. Evolution

  • Expected trajectory in 24 months (geographic expansion, product expansion, repositioning)
  • Who decides when the guide is updated, and how often
  • What would force a major revision (a rebrand, a pivot, a merger)

11. Author archetype (if SOUL.md missing)

  • Primary archetype: journalist · engineer · founder · NGO advocate · politician · consultant · executive · community lead · artist · researcher
  • Secondary archetype if hybrid
  • The principal's idiolect markers if personal-branding context (filler words, sentence length, preferred connectors, recurring metaphors, banned topics, recurring stories with allowed retelling cadence)

12. Tone of voice (if TONE.md missing)

NN/g four dimensions, position on each spectrum:

  • Funny ↔ Serious
  • Formal ↔ Casual
  • Respectful ↔ Irreverent
  • Enthusiastic ↔ Matter-of-fact

A short capture here unblocks Phase 3 codification; recommend producing a full TONE.md via samber/cc-skills@copywriting-tone-of-voice-creator afterward for a brand operating at scale.


Supporting file: references/five-layers.md

The Five Layers of Prose

Two organizing principles before codifying any layer.

Style is content, not decoration. Presentation shapes meaning — a faulty rhythm in a sentence can wreck it as surely as a wrong word. Every syntax choice, every breath point, every clause depth decision is a semantic act, not an aesthetic afterthought.

Concision is the baseline. Every word must justify its inclusion. Not brevity for its own sake — lean sentences carry more authority than padded ones because they never ask the reader to work without payoff.

Codify each layer independently. The layers are orthogonal: a brand can have distinctive lexicon and generic syntax (Liquid Death), or distinctive syntax and generic lexicon (Basecamp). Decide per layer where to conform and where to differentiate.

1. Lexicon

Word-level rules. Deterministic, testable.

1.1 Use / avoid A–Z

A 50–200 entry table. GOV.UK is the gold standard — short entries, alternatives provided, exceptions enumerated.

Term to useWhyTerm to avoidWhy avoidExceptions
customeractive subject, dignifieduserreductive outside product docs"user" OK in product docs
helpplainfacilitateempty Latinate
build / createconcreteleverageempty verbfinancial sense permitted
usedirectutilizewordy
becausecausaldue to the fact thatwordy
makeconcretedeliver (abstract)GOV.UK: "only pizzas, post and services are delivered"concrete senses OK

Why bother: banning a word without an alternative creates a vacuum writers fill with the next-worst word. Always pair ban with replacement.

1.2 Terminology table

Product names, feature names, capitalization, hyphenation, plural forms. Update on every product launch — a guide that drifts from product reality is ignored by engineers.

1.3 Jargon ladder per channel

Channel groupingSpecialist terms permitted
Long-form articlesUp to 5 per piece, with one-line glosses
Social postsUp to 2 per post, no glosses (link to glossary instead)
Email & newsletterUp to 1 per issue
Marketing copyNone unless category requires (compliance, deep-tech)

1.4 Acronyms

Spell out on first use per page (GOV.UK convention) OR allow known acronyms unexpanded (Microsoft convention). Pick one. The choice depends on audience expertise — Expert audiences resent expansion of acronyms they own.

1.5 Naming conventions

Company, products, features, methodologies. Include possessive forms. Critical for brands with non-standard capitalization (innocent, samber/lo).

1.6 Foreign words

Italics or not, accents preserved or stripped, translation in parentheses or not. Especially critical for French-origin brands writing in English. See multilingual.md.

1.7 Technical depth scale

Three-level: Layperson · Practitioner · Expert. Allocate per channel. Mixing levels within a single piece is the dominant readability failure.

2. Syntax

Sentence- and paragraph-level rules.

2.1 Sentence length distribution

Plain Language Commission default: 15–20 words mean per sentence. Category overrides in category-playbooks.md.

MetricDefault targetWhy
Mean sentence lengthcategory-specific (10–22)matches reading age, audience patience
Standard deviation≥ 6 words per 100-word windowuniform = robotic; variance = human
Sentences ≥ 25 words≤ 10%long sentences load short-term memory; rationing protects comprehension
Sentences ≤ 8 words≥ 15%punch sentences carry rhythm; their absence = monotony

Diagnose: 1- run a Python nltk.sent_tokenize + word-count script on a 1000-word sample; 2- Hemingway readability for grade-level cross-check; 3- read the piece aloud, mark every breath point.

2.2 Sentence types

  • Declarative as default
  • Rhetorical questions: max 1 per long-form piece (or banned entirely — they signal low confidence)
  • Imperatives: reserved for CTAs and how-to steps
  • Exclamations: rationed per punctuation policy (#6-punctuation-policy)

2.3 Clause depth

Max 2 levels of subordination per sentence. Beyond that, comprehension drops. Coordination preferred for B2C, subordination acceptable for B2B and industry.

2.4 Active vs passive

Active default. Passive permitted for: (a) unknown agent, (b) emphasis on object, (c) impersonal scientific register. Codify the exception list — without it, "active voice always" produces awkward science writing.

Zombie test: append "by zombies" after the verb. If it works, it's passive ("The data was analyzed [by zombies]").

2.5 Parallelism

Enforced in lists, headings, tricolons. Parallel structures must share grammatical category (all noun phrases, or all verb phrases starting with the same tense).

2.6 Paragraph length

  • Target 2–5 sentences, max 80 words on web
  • One-sentence paragraphs permitted for emphasis but ≤ 15% of paragraphs (else they lose impact)

2.7 Paragraph architecture

Codify one or two house structures. Defaults:

  • TES: Topic sentence + 2–3 evidence sentences + implication
  • PEEL: Point, Evidence, Explanation, Link
  • PAS: Problem, Agitation, Solution (newsletters, email)
  • Inverted pyramid: most important first (newsletters, news)

3. Rhythm

How sentences move together. Hardest layer to codify; easiest to audit by reading aloud.

3.1 Cadence

Target sentence-length variance such that σ ≥ 6 words per 100-word window. Lower = robotic / AI tell. See anti-patterns.md.

3.2 Breath points

One short sentence (≤ 8 words) every 3–5 sentences. Why: reading is breathing; missing breath points exhaust the reader's working memory.

3.3 Repetition

  • Anaphora (repeated openers): permitted in closings and CTAs only, or banned
  • Epistrophe (repeated endings): reserved for signature moments

3.4 Callbacks

If the hook names a person, the closing returns to that person. Why: structural symmetry is a low-cost ethos signal that the piece is built, not generated.

3.5 List patterns

  • Prefer prose-then-list (introduce in a sentence) over list-then-prose
  • Cap lists at 7 items (Miller's law)
  • Parallel grammar enforced

3.6 White space

  • Long-form: subheading every 200–350 words
  • Social posts: line break every 1–3 sentences
  • Email & newsletter: one idea per paragraph

4. Structure

Macrostructures across the piece.

4.1 Openings (hooks)

Cross-ref samber/cc-skills@copywriting-hooks for the full catalog. State 3–5 permitted hook types per channel grouping. Forbid:

  • Dictionary-definition openings ("Productivity, defined as...")
  • "In today's fast-paced world", "À l'heure du tout-numérique"
  • Self-referential openings ("This article will discuss...")

4.2 Closings

Cross-ref samber/cc-skills@copywriting-cta for end-of-article CTA codification. State 2–3 permitted closing types. Forbid: generic "Thanks for reading", "What do you think?" unless signature.

4.3 Transitions

Prefer logical connectors (because, therefore, however, but) over additive (also, moreover, furthermore). Ban "Last but not least."

4.4 Headings

Sentence case, no terminal punctuation, frontloaded with topic noun or active verb, scannable in isolation. GOV.UK rule: "Frontload your headings so that the words that match your users' tasks are at the beginning."

4.5 Subheadings

Parallel structure within an article (all questions, or all noun phrases, not mixed).

4.6 Lists

Bulleted vs numbered policy (numbered only when order matters). Leading sentence required. No nested lists beyond one level.

4.7 Asides

Parentheticals capped at one per 200 words.

4.8 Quotations

Attribution style (name, title, organization; no honorifics on second reference). Minimum quote quality bar: does the quote say something the body cannot?

4.9 Citations and links

Inline, frontloaded link text. Never "click here", "read more", "learn more".

4.10 Blockquotes

Reserved for quotes of 25+ words or for high-emphasis claims.

4.11 Reader positioning (psychic distance)

Gardner's psychic distance continuum describes how close or far the reader sits from the action, the brand, or the subject matter. In brand prose (not fiction), the spectrum runs:

DistanceRegisterExample
FarThird-person, external, historical"Acme was founded in 2005 with a mission to reduce infrastructure costs."
Medium-farCategory framing, industry truth"Most infrastructure teams spend 40% of their time firefighting, not building."
Medium-closeReader-adjacent, hypothetical"Your team is probably dealing with this right now."
CloseSecond-person present, internal experience"You open the dashboard. The number is wrong. Again."

Why it matters: distance controls emotional temperature. Far establishes authority and context. Close creates empathy and drives conversion. Uncontrolled oscillation reads as schizophrenic; deliberate oscillation creates emotional shape.

Default positions per channel

Channel groupingDefaultRationale
Long-form articlesMedium-far opening → close at key anecdote → far for analysis → close at CTAAuthority frame, then humanity, then rigor, then conversion
Social postsClose hook → medium for argument → close for CTAHook must land immediately; brevity requires intimacy
Email & newsletterMedium-close defaultPersonalization convention; inbox is a private channel
Marketing copyClose default for emotional sections, far for proof/credibilityConversion copy needs immersion; proof copy needs objectivity

Shift signals

Moving closer: second-person pronouns ("you", "your"), present tense, sensory detail, internal monologue framing ("You're wondering if…"), short sentences, concrete nouns.

Moving farther: third-person, past tense, statistics, brand history, passive constructions, abstract nouns, long sentences.

Diagnose: scan a 1500-word piece and annotate each paragraph with its distance level (F / MF / MC / C). A flat distribution (all one level) means the piece has no emotional shape. High variance without a pattern means oscillation is accidental. The target is an intentional arc.

5. Voice markers

The small set of repeatable signature elements that make the brand recognizable in a blind test. Catalogue 5–12 markers; more is unenforceable.

MarkerDefinitionExample
Signature movesRecurring rhetorical moveBasecamp's contrarian one-sentence paragraph
SignoffsFixed or templated closing lineInnocent's "Win."
Recurring metaphors2–3 metaphors the brand owns and reuses
IdiomsCurated short list (drift is a strong signal of distributed-team decay)
TaboosExplicit list of phrases the brand never uses"leverage", "delve", "in today's fast-paced world"
Intentional ticsDeliberately quirky moves (named, rationed)Innocent's lowercase brand name; Oatly's parenthetical meta-commentary

Rationing matters. Innocent's puns work because they are rationed. Unrationed quirks collapse into self-parody (the Wackywriting failure mode, per Nick Asbury).

6. Punctuation policy

Declare a position on each. The list is non-negotiable — silence creates drift.

MarkPosition
Em dashPermitted / banned (declare) — current AI-tell candidate; rationed even when permitted
En dashRanges only (2024–2026)
SemicolonPermitted in long-form, banned in social and email subject lines
ColonLists, examples, rephrased restatements; max 1 per paragraph
EllipsisBanned outside direct quotation (AI tell)
ParenthesesRationed: 1 per 200 words
ItalicsForeign words, titles of works, technical first-use emphasis
BoldScannable phrases in long-form only
Quotation marksSingle (UK) or double (US) per locale; consistent
Exclamation marks1 per 1000 words in long-form; 1 per LinkedIn post; 1 per newsletter
BracketsEditorial insertions in quotations only
HyphensCompound modifiers before nouns (well-known author); maintain a hyphenated-compound list
Oxford commaDeclare yes/no, enforce
CapitalizationSentence case for headings (Microsoft, Mailchimp, IBM Carbon default); title case for proper nouns and product names

7. Formatting policy

ElementRule
H1One per page
H2Sections
H3Sub-sections
H4+Technical docs only
Heading length≤ 70 characters
Bullets3–7 items, parallel grammar, leading sentence required, max 1 level of nesting
Numbered listsOnly when sequence matters
Code blocksLanguage tag mandatory, max 30 lines, prose explanation precedes code
ImagesSentence-case captions, descriptive alt text mandatory (WCAG 2.2)
Callouts (note, warning, tip)Rationed: 1 per 800 words in long-form
TablesOnly when relationship is two-dimensional
LinksFrontloaded link text — never "click here", "learn more", "read more"

Supporting file: references/multilingual.md

Multilingual Prose

For brands operating across languages (especially EN/FR — the dominant case for France-based operators). The core rule: one PROSE.md per language, not a translated single guide. Codify each language natively; maintain a mapping document of shared pillars and divergent rules.

Why per-language guides

A translated guide propagates the source language's rhythm into the target. French sentence structure produces long sentences with subordinate clauses; English brand prose typically favors shorter sentences. A French→English translation that preserves sentence boundaries reads as labored English. An English→French translation that preserves the original mean sentence length reads as choppy French.

The standard is to retarget, not translate. Translators must follow the target-language guide as if writing fresh.

EN variant declaration

Declare one of: US English · UK English · International English. The choice affects:

DimensionUSUKInternational
Spelling-or, -ize, color-our, -ise, colour-or, -ize (Microsoft default)
Date formatMM/DD/YYYY or "March 5, 2026"DD/MM/YYYY or "5 March 2026"YYYY-MM-DD (ISO)
Single vs double quotesDoubleSingleDouble (more globally readable)
Comma in dates"March 5, 2026""5 March 2026""2026-03-05"
Decimal separatorPeriod (3.14)Period (3.14)Period (3.14)
Thousands separatorComma (1,000)Comma or space (1,000 / 1 000)Space (1 000)
Time format12h with AM/PM24h or 12h24h

EN ↔ FR word-policing

French words permitted in English brand text

A small whitelist. Adding French outside this list usually reads as affectation:

  • Permitted (no italics, no translation): raison d'être, savoir-faire, joie de vivre, cliché, café, salon, genre, milieu, élite, fiancé(e), résumé (US) / CV (UK), déjà vu, faux pas, par excellence
  • Permitted with italics on first use: avant-garde, à la carte, en route, in vino veritas
  • Banned without translation: parcours (use "journey" or "path"), dispositif (use "system" or "framework"), enjeu (use "stake" or "issue"), accompagnement (use "support"), démarche (use "approach"), mise en œuvre (use "implementation")

English loan-words accepted in French brand text

French anglicisms cluster around marketing / tech vocabulary. Pick a position:

  • Permissive (tech, marketing, consulting brands): le marketing, le briefing, le manager, le pitch, le brainstorming, le storytelling, le timing, le mailing
  • Restrictive (cultural institutions, traditional consumer brands, government): replace with French equivalents (le brief → la note, le pitch → la présentation, le marketing → la mercatique [rarely used in practice — fallback to context-specific])
  • Banned outright in any French brand text: addressing, deliverables, leverager, prioritiser (use prioriser), implémenter (use mettre en œuvre or réaliser), supporter (use prendre en charge)

False cognates EN ↔ FR

The frequent traps for ghostwriting and translation. Reproducing the false cognate is a single-sentence credibility kill for bilingual readers.

French wordWhat writers think it meansWhat it actually means
éventuellementeventuallypossibly, perhaps
actuellementactuallycurrently, right now
importantimportantoften: large, significant in size
sensiblesensiblesensitive
déceptiondeceptiondisappointment
locationlocationrental
librairielibrarybookstore
journéejourneyday
assisterto assistto attend
supporterto supportto bear / tolerate / put up with
prétendreto pretendto claim
acheverto achieveto complete / finish
réaliserto realizeto make / produce / accomplish
consisterto consist"consister à" + verb = to involve doing
demanderto demandto ask
habithabitclothing (an outfit)
chairchairflesh
coincoincorner
painpainbread
chancechanceluck
sensible (the other way too)EN wordFR translation: rationnel, raisonnable

Syntactic transfer rules

The transfer budget corrects rhythm differences between languages.

DirectionAdjustment
FR → ENCut 20% of words. French sentences carry more subordinate clauses; English brand prose breaks them.
EN → FRPad 20% of words. French rewards more developed sentences with explicit connectors.
FR → ENReplace nominal-heavy French ("la mise en œuvre du dispositif") with verbal English ("we deploy the system").
EN → FRRestore some nominalization for register, especially in B2B and consulting.

Regionalisms and global English

Declare neutral international English when audiences are global; reserve UK or US idioms for matching audiences. Common traps:

  • US-only: "out of the gate", "ballpark figure", "touch base", "rain check"
  • UK-only: "knackered", "having a chinwag", "spot on"
  • Avoid in international: idiom-heavy sports metaphors (cricket, baseball, American football)

French regional variants

French differs across:

  • Hexagonal French (France) — default for France-based brands
  • Belgian French — septante / nonante for 70 / 90 (vs soixante-dix / quatre-vingt-dix); some lexical differences
  • Swiss French — septante, huitante, nonante; different terminology in retail / banking
  • Québécois French — significantly different vocabulary (courriel for email, magasiner for shop); different anglicism policies; English borrowings often resisted strongly
  • African French — multiple variants; significant local lexicon

Declare which variant. A French brand expanding into Québec should not assume Hexagonal French passes.

Cultural references

  • Safe references (cross-cultural): sports for athletes generally, food and seasons, holidays that are local-relevant only when audience matches
  • Forbidden for global audiences: region-specific jokes (US Super Bowl, French baccalauréat), political references, religious holidays as default context

Accessibility and inclusion

  • People-first language: "people experiencing X" over "the X"
  • Singular they (EN) — established convention since at least 1375, codified by Microsoft, Mailchimp, GOV.UK, Atlassian
  • French inclusive writing — declare a position. Three common levels:
    1. Conservative: masculine generic ("les développeurs"), historic French Academy default
    2. Inclusive parentheses: "les développeurs(euses)" — readable, contested
    3. Median point: "les développeur·euse·s" — politically loaded in France; banned in government communication since 2017; accepted in some progressive brand contexts; not all screen readers handle it gracefully
  • Bias-free language section: maintain a banned-word list for stigmatizing terms (per Microsoft, Mailchimp templates)

Translation workflow recommendation

When a brand operates in multiple languages:

  1. Author in the dominant language (usually the brand's HQ language).
  2. Translator-as-writer: never use literal translation as published content. The translator must follow the target-language PROSE.md.
  3. Terminology table is bilingual: one canonical term per language, mapped.
  4. Idiom and metaphor list per language: do not translate idioms; substitute equivalent register.
  5. False-cognate list per language pair: maintained in PROSE.md annex.
  6. Channel conventions differ across cultures: LinkedIn convention in France differs from US conventions in formality, paragraph length, first-person use. Codify per locale.

Mapping document

When multiple language guides exist, maintain PROSE-MAPPING.md documenting:

  • Shared pillars (the brand's voice principles that hold across languages)
  • Divergent rules (where each language guide departs)
  • Terminology mappings
  • Cross-references for idioms and metaphors

Supporting file: references/prose-md-template.md

PROSE.md Template

Hybrid format: narrative sections per layer + do/don't tables as an annex. The narrative teaches the why (so writers can handle edge cases); the annex provides the scannable enforcement layer.

Length target: 20–60 pages. Beyond that, writers stop reading. Below that, edge cases proliferate. Siemens reduced their brand guidelines from 2,750 to 250 pages by ruthless deletion — that is the discipline.

Skeleton

# PROSE.md — <Brand Name>

> Version <semver> · Last updated <YYYY-MM-DD> · Owner: <name / role> · Status: <draft | active | deprecated>
>
> Read alongside `TONE.md` (emotional posture) and `SOUL.md` (storyteller archetype). Visual identity lives in `DESIGN.md` and is out of scope here.

## Purpose

200 words: who this guide is for, how to use it, what it does not cover, the relationship to TONE.md and SOUL.md.

## The Prose Pillars

5–8 pillars in the form "We write X, not Y", each with a one-sentence rationale and one example. Pillars must be falsifiable. "We write short sentences with concrete subjects; we avoid abstract nominalizations" passes the test. "We write warmly" does not.

## Voice vs. Tone note

One paragraph adapting Mailchimp's formulation: "You have the same voice all the time, but your tone changes." Voice = consistent (this guide). Tone = situational (TONE.md).

## 1. Lexicon

### 1.1 Use / avoid A–Z

[Narrative paragraph explaining the lexicon's center of gravity — e.g., "Anglo-Saxon verbs over Latinate; concrete nouns over abstract; named over generic."]

[Table: 50–200 entries.]

### 1.2 Terminology

[Product names, feature names, capitalization, plural forms.]

### 1.3 Jargon ladder per channel

[Table: which specialist terms permitted in which channel grouping.]

### 1.4 Acronyms · 1.5 Naming · 1.6 Foreign words · 1.7 Technical depth scale

[As applicable; see five-layers.md for full structure.]

## 2. Syntax

[Narrative: this brand's syntax center of gravity in one paragraph.]

### 2.1 Sentence length distribution

[Mean target ± 2 words. Distribution targets. Category default reasoning.]

### 2.2 Sentence types · 2.3 Clauses · 2.4 Active/passive · 2.5 Parallelism · 2.6 Paragraph length · 2.7 Paragraph architecture

[Each subsection: rule + 1-sentence why + example.]

## 3. Rhythm

[Narrative: what cadence sounds like read aloud.]

### 3.1–3.6 Cadence · Breath points · Repetition · Callbacks · List patterns · White space

## 4. Structure

### 4.1 Openings

[3–5 permitted hook types with example openings from prior brand corpus.] [Forbidden openings list.]

### 4.2 Closings · 4.3 Transitions · 4.4 Headings · 4.5 Subheadings · 4.6 Lists · 4.7 Asides · 4.8 Quotations · 4.9 Citations · 4.10 Blockquotes

## 5. Voice Markers

[5–12 markers with rules of use and rationing.]

### 5.1 Signature moves · 5.2 Signoffs · 5.3 Recurring metaphors · 5.4 Idioms · 5.5 Taboos · 5.6 Intentional tics

## 6. Punctuation Policy

[The full table from five-layers.md, adapted to this brand's positions.]

## 7. Formatting Policy

[Heading hierarchy, lists, code blocks, images, callouts, tables, links.]

## 8. Channel Overrides

[One section per in-scope grouping. Each section: deltas on sentence length, paragraph length, hook types, closing types, formatting, CTA.]

### 8.1 Long-form articles

### 8.2 Social posts

### 8.3 Email & newsletter

### 8.4 Marketing copy

## 9. Cultural & Linguistic Adaptation

[English variant (US/UK/intl); French↔English handling; false cognates; transfer budgets; accessibility/inclusion.]

## 10. Anti-LLM Countermeasures

[Banned lexical tells, structural tells, punctuation defaults. The rules LLMs do not follow by default — that is the durable defense.]

## 11. Sample Bank

### 11.1 Before/after pairs (≥ 10)

For each:

- Rule violated
- Original
- Rewrite
- Rule applied
- Why the rewrite is better (one sentence)

### 11.2 Exemplar pieces (≥ 3, annotated paragraph by paragraph)

### 11.3 Anti-exemplars (≥ 2, de-identified)

### 11.4 Hook bank (30+ approved openings)

### 11.5 Closing bank (15+ approved closings)

### 11.6 Transition bank (20+ approved transition phrases)

## 12. Ghostwriting Addendum (per principal, if applicable)

[Per-principal idiolect: 10 signature openings, 5 signature closings, 3 recurring stories with allowed retelling cadence, list of banned topics, list of preferred connectors, average post length, line-break convention.]

---

## Annex A — Do/Don't Quick Reference

The scannable layer. One table per layer; writers can audit a draft in 10 minutes.

### A.1 Lexicon

| ✅ Do | ❌ Don't |
| --- | --- |
| Use plain Anglo-Saxon verbs | Use empty Latinate verbs (leverage, facilitate, utilize) |
| Spell out acronyms on first use | Assume all readers know the acronym |
| Use product names exactly as registered | Improvise capitalization |

### A.2 Syntax

| ✅ Do | ❌ Don't |
| --- | --- |
| Vary sentence length (σ ≥ 6 words per 100-word window) | Write uniformly long or uniformly short sentences |
| Active voice as default | Use passive without a documented reason |
| Cap subordination at 2 levels | Stack subordinate clauses |

### A.3 Rhythm

| ✅ Do | ❌ Don't |
| --- | --- |
| Place a breath sentence (≤ 8 words) every 3–5 sentences | Forget to breathe |
| Use parallel structure in lists and tricolons | Mix grammatical categories in a list |
| Cap lists at 7 items | Write 12-item lists with no grouping |

### A.4 Structure

| ✅ Do | ❌ Don't |
| --- | --- |
| Front-load headings with the topic noun | Open with throat-clearing ("Introduction to...") |
| Use logical connectors (because, therefore, however) | Use additive filler (also, moreover, furthermore, last but not least) |
| Frontloaded link text | "click here", "learn more", "read more" |

### A.5 Voice markers

| ✅ Do | ❌ Don't |
| --- | --- |
| Deploy signature moves at the declared rate | Overuse a signature move into self-parody |
| Maintain the taboo list | Drift into category-default phrasings |

### A.6 Punctuation

| ✅ Do | ❌ Don't |
| --- | --- |
| Enforce the Oxford comma decision consistently | Switch within a piece |
| Ration exclamation marks per the policy | Use exclamations as enthusiasm performance |
| Banned em dash → use comma, colon, parens, or period (if banned) | Keep em dashes when policy says no |

### A.7 Channel overrides

| Channel | Mean sentence length | Paragraph length | Hook style | CTA |
| --- | --- | --- | --- | --- |
| Long-form | 14–18 | 2–5 sentences | Scene · contrarian · stat · concrete detail | Practical next step / callback |
| Social | 8–12 | 1–3 sentences | Bold claim · direct problem · concrete detail | Specific reply prompt |
| Email | 10–14 | 1–2 sentences | Personal frame · curiosity gap | Single primary CTA in P.S. |
| Marketing copy | 8–12 | 1–3 sentences | Promise · direct problem · authority | Direct action button |

---

## Changelog

| Date       | Version | Change        | Author |
| ---------- | ------- | ------------- | ------ |
| YYYY-MM-DD | 1.0.0   | Initial guide | name   |

Notes on populating the template

  • Pillars are mandatory. Without falsifiable pillars, writers default to invented rules.
  • Sample bank is the most-read section. Lead with it in onboarding; treat it as the front door, not the appendix.
  • Annex tables are co-located by layer. Editors read top-down narrative for understanding; writers spot-check via the annex on every piece.
  • The changelog is part of the trust. A guide updated visibly is a guide writers trust.

How do I install Copywriting prose creator in Cursor, Claude Code, or Codex?

Run npx skills add samber/cc-skills --skill copywriting-prose-creator in the project where you want it, then ask your agent for the skill by name. The --skill flag installs only Copywriting prose creator, not every skill in the repository.

Where does Copywriting prose creator come from and what license is it under?

Copywriting prose creator comes from the samber/cc-skills repository on GitHub. That repository has 152 GitHub stars. The skill is published under the MIT license.

Prefer plain text? Read the Copywriting prose creator guide as markdown.