Mind to Mind

Markdown's Tower of Babel

Markdown’s Tower of Babel

One name, many grammars: the standard that never was

Mention “Markdown,” and most developers think of a single, universal standard. It isn’t. The name actually covers a messy family of related, yet mutually incompatible grammars. Back in 2004, John Gruber released the original version as a basic Perl script accompanied by a prose description. Neither specified the syntax with anywhere near the precision needed for independent parsers to agree. Ten years later, CommonMark (2014) finally brought order to the core syntax. But it only goes so far. Real-world documents need tables, footnotes, math, metadata, task lists, and callouts which are all features that CommonMark leaves entirely to a wild west of custom extensions. How hosts support these extensions varies widely.

This leaves us with a problem: a standard .md file has absolutely no way to declare which specific grammar it was written for. Feed the exact same bytes into two different renderers, and you will get two different renderings. This article catalogues these major points of divergence, complete with references to the primary specs.

The original description was not a specification

Gruber’s 2004 syntax page doesn’t specify behavior so much as illustrate it by example, deferring hard edge cases to Markdown.pl. It states several rules loosely, if at all. For instance, list continuation lines can be indented or not (“if you want to be lazy, you don’t have to”). Want a hard line break? Put “two or more spaces” at the end of a line. What about intraword emphasis, nested-list indentation depth, or when an HTML block actually ends? The original document simply ignores them. To make matters worse, Markdown.pl itself was abandoned after December 2004. As the CommonMark project later noted, it “was quite buggy, and gave manifestly bad results in many cases,” making bug-ward compatibility a hopeless goal.

Naturally, independent implementers filled these gaps as they saw fit. Take Babelmark (started by Fortin, passed to MacFarlane, now maintained by Mutel), which renders a single snippet through dozens of processors side-by-side. Try feeding it a short input with two-space list indents, intraword asterisks, blank lines between list items, or a # header missing its following space. The processors split into wildly different camps. As the CommonMark project succinctly put it: “Because there is no unambiguous spec, implementations have diverged considerably over the last 10 years.”

CommonMark specifies the core and stops there

In September 2014, John MacFarlane (of Pandoc fame) and Jeff Atwood (co-founder of Stack Overflow and Discourse) teamed up with developers from GitHub, Reddit, and Meteor to launch CommonMark. They delivered a highly detailed, normative specification, a conformance suite containing hundreds of test cases, and robust reference implementations in C (cmark) and JavaScript (commonmark.js). As of today, the current release is 0.31.2, published in January 2024.

Initially, the creators boldly named it “Standard Markdown.” John Gruber objected almost immediately on launch day. Jeff Atwood’s fascinating postmortem details the drama and the frantic requests to rename the project, hand over the domain, and apologize, leading to the team rejecting " Common Markdown" for a name completely stripped of Gruber’s trademarked word. Given that Gruber never endorsed the final specification, the word “Markdown” on its own still doesn’t point to any single, official syntax.

Within its narrow boundaries, CommonMark’s precision is surgical. Tabs expand to 4-space stops in block contexts. A list item’s content must align with the marker width plus one to four spaces. Indented code requires exactly four spaces, and hard breaks need two spaces or a backslash. It even defines seventeen incredibly dense delimiter-run rules for emphasis, and details seven distinct start and end conditions for raw HTML blocks.

But that precision comes at a cost: scope. CommonMark covers only the absolute basics such as headings, paragraphs, block quotes, basic lists, raw HTML, and code blocks. It completely lacks support for tables, footnotes, task lists, strikethrough, definition lists, math notation, element attributes, and metadata blocks. Every single one of these modern necessities is delegated to custom extensions. Worse, not a single extension has been standardized in the decade since.

Even John MacFarlane has grown frustrated with the limitations. In his critique, Beyond Markdown, he admits that the emphasis rules “still leave cases undecided,” reference links can’t be parsed until the end of a document, and list-item complexity exists mostly to support legacy indented code blocks. Instead of trying to fix these deep flaws in CommonMark, MacFarlane designed an entirely new markup language called Djot.

The extension zoo

Rather than uniting, major platforms have carved out their own custom territories. If you hear a file is written in " Markdown," that tells you almost nothing until you know which specific dialect from the following lineup is in play:

Dialect Home Adds beyond CommonMark Notes
PHP Markdown Extra (2005) Michel Fortin Pipe tables, footnotes, definition lists, abbreviations, fenced code (~~~), ID and class header attributes, Markdown parsing within HTML blocks Served as the blueprint for almost all subsequent table and footnote syntaxes.
MultiMarkdown (2005) Fletcher Penney Tables, footnotes, academic citations, metadata headers, math blocks, cross-referencing Marketed as a “superset of the Markdown syntax”; grew out of a heavily patched version of the original Perl script.
Pandoc Markdown John MacFarlane Four distinct table types (simple, multiline, grid, pipe), footnotes, definition lists, bibliographies/citations, LaTeX math, fenced divs, bracketed spans, custom attributes, YAML front matter High-powered but highly non-portable. Users can toggle every single extension individually using a boolean-like compiler flag style (+ext/-ext).
GitHub Flavored Markdown (spec 0.29-gfm, 2019) GitHub Pipe tables, task lists, strikethrough, extended autolink detection, HTML element filtering Officially “a strict superset of CommonMark.” However, GitHub.com quietly adds extra features like repository footnotes, TeX math, alert callouts, and Mermaid diagrams that don’t actually exist in the formal GFM spec.
kramdown Thomas Leitner, default in Jekyll Tables, footnotes, definition lists, inline abbreviations, attribute lists, math blocks, block comments The default Jekyll parser. Many blog authors write GFM-style Markdown and get caught off guard by kramdown’s strict layout requirements.
markdown-it Vitaly Puzrin, Alex Kocharin Modular plugin system extending CommonMark: footnotes, definition lists, text marking, superscript, subscript, custom containers, and emoji shortcuts. The engine behind VS Code’s markdown preview, VuePress, and VitePress. Out-of-the-box behavior is highly unpredictable since it depends entirely on active plugins.
Obsidian Flavored Markdown Obsidian Internal wikilinks, file embeds, block identifier anchors, syntax comments, text highlighting, callout banners Promises compatibility with GFM, but heavy reliance on its proprietary syntax makes files almost impossible to render correctly outside Obsidian.
Slack mrkdwn Slack Bold, italic, and strikethrough markup, custom link syntax; explicitly lacks support for headings, tables, or image embedding. Not a true Markdown dialect. As Slack’s developers admit, it’s merely “inspired by markdown, but uses different rules.”
Discord Discord Highly restricted subset: bold, underline formatting, spoiler tags, subtext, shallow headings, basic lists. Reassigns underscores to mean underlines instead of emphasis. Standard markdown structures like tables and inline images are completely absent.
Reddit Reddit Old Reddit relies on a custom fork featuring spoiler tags, superscript, and basic tables. New Reddit leverages a custom rich-text editor. The platform hosts two different rendering engines simultaneously, and they regularly disagree on formatting.
MDX mdx-js project Embedded JSX components, Javascript import/export statements, dynamic curly-brace expressions. A major trap for copy-pasted text. Curly braces and angle brackets that render as literal characters elsewhere will break the parser here.

What does this fragmentation teach us? Three key takeaways jump out immediately:

Where the lines are drawn

When you move a Markdown file from one renderer to another, these specific syntax elements are the most likely to break your layouts.

Hard line breaks. How do you force a line break? Gruber’s original spec requires two trailing spaces, a rule that modern text editors and aggressive linting tools love to strip away automatically. To solve this, CommonMark allows a trailing backslash. But platforms like GitHub complicate things further: they treat a bare newline as a soft break in repository .md files, but force it as a hard break inside issues and pull requests. Because Pandoc makes this behavior toggleable via its hard_line_breaks extension, you can never quite be sure how your paragraphs will wrap.

List indentation. The old Markdown.pl script required a clean four-space indent for nesting and continuation. CommonMark, on the other hand, dynamically aligns content with the first character after the marker (meaning column 2 for - and column 3 for 1. ), using tab stops set to 4. Writing text for one set of rules completely breaks nesting in the other. To add to the confusion, kramdown flatly bans mixing different bullet markers within a single list, even though Gruber’s original guide explicitly allowed it.

Tables. Let’s start with a simple fact: tables do not exist in core CommonMark. While the pipe syntax popularized by PHP Markdown Extra is widely used, implementations disagree on nearly everything else. Must you include outer pipes? Is a header row mandatory? Can a single cell contain block-level content like lists? How do you escape a literal pipe inside a code span? What happens when a row has the wrong number of columns? While Pandoc offers simple, multiline, and grid layouts, other parsers will treat them as plain, unformatted text.

Footnotes. While the [^1] syntax is incredibly common, it doesn’t actually exist in the CommonMark or GFM specs. GitHub.com supports it in repository files, but completely ignores it in wiki pages. Even when supported, different engines generate backlinks differently and disagree on how to format multi-paragraph notes.

Math rendering. Attempting to render LaTeX-style math using $ symbols is a recipe for trouble when writing about currency. GitHub handles this via $…$, $$ blocks, and special ```math fences. But kramdown forces $$ for both inline and display math, while Pandoc enforces strict whitespace requirements next to the dollar signs. Obsidian, Jupyter, and Typora all follow their own custom rules.

Front matter. That block of YAML metadata bounded by --- lines is the lifeblood of static site generators, but it is entirely foreign to Markdown. The CommonMark maintainers have made their stance clear: front matter is out of scope. They advise apps to strip it out before feeding the rest to the parser. If you don’t, an unaware compiler will treat the opening --- as a horizontal rule, and use the closing --- to turn your metadata lines into a massive header.

Admonitions. Want to show a warning box? You have to choose between four entirely separate syntaxes. GitHub uses blockquotes with specific keywords like > [!NOTE]. Obsidian uses a similar syntax but supports custom classes and folding. Python-Markdown prefers the indented !!! note format, while Pandoc relies on fenced divs like ::: {.callout-note}. Feed the wrong syntax to a parser, and your beautiful callout degrades into a messy pile of brackets and colons.

Underscore emphasis. Can you use underscores mid-word? CommonMark ignores intraword underscores so that snake_case_variables don’t break. Gruber’s original script and older parsers, however, will italicize the middle of your variable name. Meanwhile, Slack ignores asterisks for emphasis entirely, using single underscores instead.

Raw HTML. CommonMark defines seven complex block types to govern where HTML begins and ends. But GFM’s security features strip out elements like <script>, <title>, and <iframe>. While PHP Markdown Extra and Pandoc can actively parse Markdown formatting hidden inside HTML tags, CommonMark will ignore it completely.

Heading anchors. If you want to link directly to a section like #installation, you are at the mercy of the renderer’s slug generation algorithm. How spaces, uppercase letters, punctuation, and duplicate headings are converted into IDs varies so much across GitLab, GitHub, and Pandoc that your internal links are almost guaranteed to break somewhere.

Wikilinks. The handy [[Page Name]] syntax is loved by Obsidian, Logseq, and Roam users. However, it is entirely absent from any formal Markdown specification. Outside of those specific personal knowledge base tools, they simply will not resolve.

Why this fragmentation is dangerous

We’ve collectively forced Markdown to act as a universal document interchange format. It was never built for this. Today, documentation pipelines convert it to HTML, PDF, and man pages; note-taking apps import and export it; databases store it only for different services to render it later using completely different engines. Every single transition is a gamble. Worse, unlike XML, Markdown is designed never to fail. It will always output something, even if that something is a mangled, unreadable mess.

Our tools can’t save us because Markdown files cannot declare their own dialect. An .md file has no DOCTYPE, no schema reference, and no media-type parameter. Editors and linters are forced to guess. Run a formatter like Prettier, and it might silently “normalize” list indents to its own preferred standard, completely altering how your lists nest when parsed by a stricter engine.

This issue is compounding rapidly thanks to AI. Large language models, trained indiscriminately on every flavor of Markdown, regularly hallucinate a hybrid syntax. They will confidently mix GFM tables, MultiMarkdown footnotes, LaTeX math, and Obsidian-style callouts in a single response. This chaotic output is then copy-pasted into wikis, fed into build pipelines, and committed to repositories. AI didn’t create this fragmentation, but it is accelerating and distributing it at an unprecedented scale.

The core problem is a fundamental mismatch of purpose. We are treating a format designed for casual, readable authoring as if it were a rigid document interchange standard. John Gruber explicitly warned against this from the start: “HTML is a publishing format; Markdown is a writing format.” The only serious attempt to specify it came a decade too late, covered only a fraction of the syntax in use, and was barred from even using the name “Markdown.”

How to survive the Babel

Don’t hold your breath waiting for a unified standard. The CommonMark committee has explicitly ruled out front matter and moves at a snail’s pace on extensions. Individual platforms have zero incentive to align, and John Gruber has checked out of the conversation entirely. That leaves us with only one viable strategy: defensive scoping.

If you are an author: Find out exactly what renderer your platform uses and stick strictly to its documented rules. If your files need to be highly portable, limit yourself to the absolute lowest common denominator: CommonMark plus GFM tables, task lists, and strikethrough. Use a backslash to force hard breaks. Always indent your nested bullet lists by exactly four spaces—this satisfies both CommonMark’s complex alignment logic and older, legacy parsers. Finally, treat footnotes, math blocks, callouts, and front matter as highly fragile, non-portable extensions.

If you are an implementer: Build your tools on top of a strictly CommonMark-compliant parser and let users explicitly toggle extensions by name (similar to Pandoc’s +extension flag). Publish your conformance scores against the official CommonMark test suite and Babelmark. If you are building an editor, allow users to select their target dialect and lint their files against it. When generating Markdown automatically, whether via templates or AI, always target the smallest possible dialect. When in doubt, fall back to raw HTML; it is the one construct every parser is forced to pass through.

If you are starting a new project: Ask yourself if you actually need Markdown at all. John MacFarlane’s Djot resolves almost all of these issues in a single, clean specification that includes tables, footnotes, math, attributes, and generic containers, all while being significantly easier to parse. Alternatively, AsciiDoc and reStructuredText have offered stable, well-defined extension systems for twenty years. They might not be as fashionable, but they guarantee that the exact same input will render the same way every single time.

Primary Sources & Reading

Tags: