The Hidden Cost of Markdown Word Inflation
When you are writing technical documentation, a blog post, or a newsletter in Markdown, getting an accurate word count is surprisingly difficult. Standard word counters are designed for plain text. When you paste Markdown into them, they blindly count syntax tokens as actual words.
This causes three main issues for writers and engineers:
- Inflated Reading Times: Code snippets, which are often scanned rather than read linearly, artificially boost the estimated reading time of an article.
- Skewed Publisher Limits: If a publisher requests a strict 1,500-word limit, code blocks, inline code and structural markers can push a raw count well over the real prose count.
- Inaccurate Translation Pricing: If you are paying for localization by the word, paying a translation service to parse JSON payloads inside a code block is a waste of budget.
To get precise prose metrics, you must isolate the human-readable text from the structural CommonMark markup. A common misconception is that link URLs add extra words. In fact, inline link URLs add no extra tokens to wc -w because they contain no spaces (5 vs 5 words). What actually inflates counts: fenced code, inline code, and standalone markers (#, -, >). On a sample document, wc -w counts 33 words, while there are only 13 prose words.
Method 1: The Command Line Approach (macOS & Linux)
If you are working in a terminal environment, you can combine standard Unix utilities like sed and wc to strip out large fenced code blocks before counting the remaining words.
The Quick and Dirty sed command
Here is a quick and dirty bash pipeline that removes any lines that start with three backticks and then passes the result to the POSIX wc utility:
sed '/^```/,/^```/d' document.md | wc -w
Limitations of the quick sed approach:
This simple sed range only works with plain triple-backtick fences. It misses ~~~ fences, indented fences, and leaks the content of nested 4-backtick fences. On our sample document, this tested result returns 24.
The robust awk and sed pipeline
For a more accurate count that correctly handles CommonMark spec: fenced code blocks, you can use this tested awk and sed pipeline:
cat > strip-fences.awk <<'EOF'
{
line = $0; sub(/^ {0,3}/, "", line)
if (match(line, /^(`{3,}|~{3,})/)) {
f = substr(line, 1, RLENGTH)
if (fence == "") { fence = f; next }
if (substr(f, 1, 1) == substr(fence, 1, 1) && length(f) >= length(fence) && line ~ /^[`~]+[ \t]*$/) { fence = ""; next }
}
if (fence == "") print
}
EOF
awk -f strip-fences.awk document.md \
| sed -E 's/`[^`]*`//g; s/!?\[([^]]*)\]\([^)]*\)/\1/g; s/^[[:space:]]*(#+|>+|[-*+]|[0-9]+\.)[[:space:]]+//' \
| wc -w
Tested output: 13 (vs wc -w 33 and the quick sed 24).
Method 2: The AST Parser via Node.js
The most robust way to calculate accurate metrics is to use a parser that actually understands the Abstract Syntax Tree (AST) of Markdown. Here is a tested Node/mdast version that walks the Markdown syntax tree and counts only text nodes.
First, install the dependencies:
npm i mdast-util-from-markdown@2 unist-util-visit@5
Then, run this Node script:
import { readFileSync } from 'node:fs';
import { fromMarkdown } from 'mdast-util-from-markdown';
import { visit } from 'unist-util-visit';
const tree = fromMarkdown(readFileSync(process.argv[2], 'utf8'));
let text = '';
visit(tree, (node) => {
if (node.type === 'code' || node.type === 'inlineCode' || node.type === 'html') return 'skip';
if (node.type === 'text') text += node.value + ' ';
});
console.log(text.split(/\s+/).filter(Boolean).length);
Tested output: 13. This relies on the mdast-util-from-markdown library to build the tree.
Quick Word and Reading-Time Stats
For quick statistics, Utiliome provides a Markdown Word Count tool. It strips link URLs and formatting characters (#, *, `, _, ~, >, +, -, =, |) as well as HTML tags to report words, characters, paragraphs, headings, and reading time (at 200 words per minute). Note that this tool uses regex rather than an AST, and it does count code blocks. For code-free counts, use the scripts above. The tool runs in your browser, so the file is not uploaded.
3-Point Pre-Flight Checklist for Accurate Markdown Metrics
Before submitting your next technical article, follow this pre-flight check to ensure your word count is accurate and your workflow is secure:
- Exclude Fenced Code Blocks: Ensure your counting tool ignores fenced code blocks if you only want prose metrics.
- Strip URL Destinations: Verify that the destination links are not inflating character counts.
- Compare Methods: Run the quick sed count and the awk pipeline on the same file; a large gap means fences the simple command missed.
