John Gruber Left the Edge Cases Undefined in 2004. The Spec Wars That Followed Lasted a Decade.
In 2004, John Gruber published Markdown on his blog Daring Fireball, releasing a text-to-HTML conversion tool and a set of plain-text formatting conventions to accompany it. Aaron Swartz, who collaborated on the conceptual design, is sometimes credited as a co-creator, though Gruber later disputed this characterization. The accompanying Perl script was fewer than 1,500 lines. The specification, a single web page describing the syntax, deliberately left dozens of edge cases unresolved. Gruber's stated goal was a writing format, not a programming language. How it handled ambiguous nesting, conflicting list markers, or nested emphasis within links was, in his view, a matter of implementation preference rather than something that required formal standardization.
The consequence of that decision was a decade of incompatible parsers, a naming dispute that involved accusations of trademark violation, and eventually a nearly 100-page formal specification to resolve the ambiguities that Gruber had considered unimportant.
Gruber's core insight was that plain-text email conventions already contained an informal Markdown. People writing email in the 1990s and early 2000s were already using asterisks around words to indicate emphasis, using hash marks or equals signs to indicate headings, and using hyphens to create lists. These conventions emerged organically because they made sense visually in monospace text. Gruber codified those conventions into a formal syntax and wrote a Perl script that converted them to valid HTML. The output was not a simplification of HTML. It was a human-readable source format that produced HTML as output, keeping the two representations separate.
How Markdown Spread and Why It Fragmented
Markdown spread because it solved a practical problem that was genuinely irritating to a large population of writers. Technical bloggers and developers who wrote content for the web did not want to type HTML tags by hand, but they needed the output to be valid HTML. Markdown let them write in a format that was readable as plain text and that produced correct HTML automatically.
Jeff Atwood and Joel Spolsky adopted Markdown for Stack Overflow when the site launched in 2008, exposing the syntax to a large developer audience and cementing it as the standard for developer-oriented writing contexts. GitHub adopted Markdown for repository README files, issue comments, pull request descriptions, and wiki pages. The convention spread through the developer community as GitHub grew, until Markdown became the near-universal default for developer documentation and technical writing.
But implementations multiplied without coordination. Developers who needed Markdown support in their languages wrote parsers in Python, Ruby, JavaScript, PHP, Java, and dozens of others. Because the original specification had gaps, each parser resolved ambiguous cases differently. A document that rendered one way in Gruber's original Markdown.pl rendered differently in Python-Markdown, kramdown, or Redcarpet. GitHub added its own extensions for task list checkboxes, code fences with language identifiers, automatic URL linking, and strikethrough text, and created GitHub Flavored Markdown as a formal variant. MultiMarkdown added metadata blocks, footnotes, and table support. Pandoc's Markdown extended further.
By 2012, the fragmentation was severe enough that Jeff Atwood and John MacFarlane, who had created the pandoc document conversion tool, began working on a formal specification to resolve the ambiguities. They initially called it Standard Markdown, a name that prompted Gruber to object publicly. The project was renamed to CommonMark in 2014. The CommonMark specification is nearly 100 pages long and covers parser behavior for hundreds of edge cases through a formal test suite that any conforming implementation must pass.
The Parsing Challenge
The process of converting Markdown to HTML is conceptually simple: a parser reads the Markdown source and builds a document tree representing the content structure, then a renderer emits HTML elements for each node. The complexity is in the parsing itself.
Markdown uses two different parsing modes simultaneously. Block-level elements, including paragraphs, headings, code blocks, blockquotes, and lists, are determined by line-level analysis. The structure of a document is determined by which lines begin with specific characters and how those lines relate to surrounding lines. Inline elements, including emphasis, strong emphasis, links, images, and inline code, are determined by character-level parsing within the text of each block element.
These two levels interact in ways that require careful rule prioritization. A document might contain a list item whose text includes an emphasis span that contains a backtick that might or might not be code. The ambiguity is resolved by a specific priority ordering: code spans are parsed first, which prevents backtick characters from being misinterpreted as parts of other constructs. Getting this ordering wrong produces incorrect output.
CommonMark's specification resolves these priority questions explicitly. That explicitness is what makes CommonMark implementations interoperable. Any parser that passes the CommonMark test suite will produce the same output for ambiguous inputs as any other conforming parser.
Markdown Flavors and Their Differences
The most widely encountered Markdown flavors differ in specific areas that matter for practical use.
GitHub Flavored Markdown adds tables, task list checkboxes ([ ] and [x] syntax), strikethrough text using double tildes, and automatic linkification of URLs without requiring explicit bracket syntax. These features are ubiquitous in GitHub issues and pull requests. A document that uses GFM-specific syntax converted through a CommonMark-only parser will produce incorrect output or literal text where a table or checkbox should appear.
MultiMarkdown adds footnotes using [^identifier] syntax, definition lists, metadata blocks at the start of the document, and cross-references. These features are used in academic and long-form writing contexts where footnotes and citations matter.
Pandoc's Markdown is the most extended flavor, adding support for LaTeX math notation, raw LaTeX and HTML blocks, extension-based configuration, and extensive output format control. Pandoc can convert Markdown to PDF, Word, EPUB, LaTeX, and dozens of other formats, which is why it became standard in academic and publishing workflows.
The practical consequence: when choosing a Markdown-to-HTML converter for a specific use case, verify whether it supports the flavor of Markdown you are using. Copy-pasting a GitHub README into a CommonMark converter may not handle the GFM table syntax. Pasting a Pandoc-flavored academic document into a standard parser will produce garbled footnote markers.
Why Output Quality Matters
The semantic quality of the HTML output from a Markdown converter matters for practical reasons beyond appearance.
Heading hierarchy is the most consequential. Correct heading structure, from h1 through h6 without skipping levels, is required for accessibility. Screen readers use the heading hierarchy to navigate documents. A converter that emits h1 followed by h4 produces technically valid HTML that violates accessibility guidelines and confuses navigation. WCAG 2.1 Success Criterion 1.3.1 requires that information and relationships conveyed through presentation are programmatically determinable, which heading hierarchy is.
Heading structure also affects how search engines interpret page content. Search crawlers use heading tags to understand document structure and the relative importance of content sections. A document where all headings are h2 regardless of their semantic level loses the structural signal that well-nested headings provide.
Link construction is another area where converter quality varies. Markdown allows relative links that are interpreted relative to the document's location. A converter that emits these links without resolving the base URL may produce links that work when the converted HTML is in the original location but break when moved.
Code block handling differs between converters in whether they escape HTML entities within code blocks. Code that contains angle brackets or ampersands must have those characters escaped in HTML output to prevent the browser from interpreting them as markup. A converter that fails to escape these characters produces output that looks correct in many cases but breaks when code contains characters that HTML interprets as markup tokens.
Conclusion
ToolHQ's Markdown to HTML converter uses a standards-compliant parser and shows the rendered output alongside the HTML source so you can verify the conversion before using it. The tool handles CommonMark syntax and produces well-structured, properly escaped HTML output. For flavors that extend CommonMark with tables or task lists, the relevant extension syntax is supported.
The spec wars that followed Gruber's 2004 publication resolved into a workable equilibrium: CommonMark as the interoperable base, GitHub Flavored Markdown as the dominant extended flavor for developer content, and Pandoc as the tool for conversion-heavy workflows. None of these would exist in their current form if Gruber had specified the edge cases that he chose to leave undefined in 2004. The fragmentation produced competition, and the competition produced better tools than a single canonical implementation would have.
Frequently Asked Questions
What is CommonMark and why was it created?
CommonMark is a 2014 specification that resolves ambiguities in John Gruber's original 2004 Markdown spec. It includes a 100-page formal spec and test suite to ensure consistent behavior across parsers.
What is the difference between Markdown flavors?
Flavors like GitHub Flavored Markdown, MultiMarkdown, and Pandoc Markdown extend the original spec with features like tables, footnotes, task lists, and code fences. They are not interchangeable.
Why does Markdown to HTML conversion quality matter?
Poorly structured HTML output can skip heading levels, malform links, or produce inaccessible markup that affects screen reader behavior and SEO semantic structure.