Benchmark
We saved six reports with 5 editor engines. All 5 rewrote them.
Load a markdown file. Save it without touching anything. Count the lines that changed. On every serialising engine we tested, the answer was never zero.
Method
Six report files were built to look like the markdown people actually receive, each seeded with the constructs that break editors:
| File | Simulates | Byte-level traps |
|---|---|---|
| 01 Pentest findings | External pentest report: YAML frontmatter, severity tables with alignment, CVSS, HTTP and JSON proofs of concept, task lists, GitHub alerts, details blocks, footnotes, escaped pipes | LF |
| 02 AI code audit | Agent security review: XML finding tags, emoji severity table, diff blocks, nested quotes, CJK, Greek, Hebrew and Arabic text | CRLF + UTF-8 BOM |
| 03 SOC 2 readiness | 8-column control matrix, exceptions, nested lists with tabs, evidence image with title, trailing spaces | Mixed CRLF and LF, tabs, trailing whitespace |
| 04 Architecture spec | Design RFC: TOML frontmatter, Mermaid, inline and block math, a definition list, reference-style image, HTML reviewer comment | LF |
| 05 LLM chat export | Chat transcript: nginx configs, tables, Japanese, Chinese, Korean, right-to-left text, literal escape sequences | LF |
| 06 Edge cases | Setext headings, every list marker, odd ordered lists, lazy continuation, entities, nested fences, pipe-less tables, three horizontal-rule styles | No final newline |
Each file was loaded and saved with no edits by seven configurations (five serializing engines plus CodeMirror 6 saved two ways), using each engine's documented load and save calls: remark (unified 11 with GFM, frontmatter and math), prosemirror-markdown 1.13, marked to HTML to turndown 7, Milkdown 7.22 (commonmark and GFM presets), Tiptap 3.31 with the official markdown extension, CodeMirror 6 with a naive doc.toString() save, and CodeMirror 6 with the patch-on-save reference technique. These are engines, not the branded apps. Closed GUI apps such as Typora, and Obsidian and VS Code's rendered editors, were not tested directly; Obsidian is built on CodeMirror 6, so it is closer to the CodeMirror rows.
The measure is the percentage of original lines that a line diff marks as removed. 0 means byte-identical. The corpus is pinned by SHA-256 and every package version is pinned in the lockfile.
Result A: save with no edits
| Report | remark | prosemirror-markdown | marked + turndown | Milkdown 7 | Tiptap 3 | CodeMirror 6 (naive save) | Patch-on-save (reference) |
|---|---|---|---|---|---|---|---|
| 01 Pentest findings (LF) | 32.5% | 45.4% | 44.8% | 36.2% | 38.7% | 0 | 0 |
| 02 AI code audit (CRLF + BOM) | 88.7% | 100% | 100% | 88.7% | 100% | 100% | 0 |
| 03 SOC 2 readiness (mixed CRLF/LF) | 64.8% | 64.8% | 63.4% | 64.8% | 59.2% | 33.8% | 0 |
| 04 Architecture spec | 14.3% | 37.4% | 38.5% | 18.7% | 25.3% | 0 | 0 |
| 05 LLM chat export | 11.2% | 7.1% | 16.3% | 11.2% | 10.2% | 0 | 0 |
| 06 Edge cases (no final newline) | 36.0% | 42.1% | 48.2% | 38.6% | 40.4% | 0 | 0 |
| Byte-identical files | 0/6 | 0/6 | 0/6 | 0/6 | 0/6 | 4/6 | 6/6 |
Milkdown and Tiptap ran with default presets, without frontmatter, math, footnote or alert extensions. Products built on them often add such plugins, which would fix some construct-level losses but not the reformatting, escaping and line-ending changes that come from re-serialising. The last column is a reference implementation of patch-on-save, the technique AsItIs uses, not the AsItIs app itself.
What they broke
Content lost
- Tiptap turned the SOC 2 evidence image and the spec's reference-style image into plain text, removed the
<details>wrapper, the<dl>definitions and the HTML reviewer comment, and turned the YAML frontmatter into a heading. - Milkdown (default presets) dropped the reference-style image entirely and logged an internal error doing it. It turned the opening
---of the YAML frontmatter into a horizontal rule, so the metadata became body text, and stripped the BOM. - marked + turndown removed all four
<finding>tags with their severity attributes from the AI audit, so the severity metadata is gone, along with<details>,<dl>,kbd,mark,supand the reviewer comment. Task items became plain bullets. - prosemirror-markdown has no table node in its default schema: every table in the pentest report and the SOC 2 control matrix was flattened into one paragraph of pipes. It also escaped footnotes and merged their definitions, escaped task checkboxes, and broke both frontmatter blocks and the math block.
- remark kept most constructs, with frontmatter, GFM and math plugins enabled, but still changed 11% to 89% of lines.
Meaning silently changed
- Every serialising engine escaped the GitHub alert:
> [!CAUTION]became> \[!CAUTION], so GitHub renders an ordinary quote instead of the red Caution box. In a security report, that box is the handling notice. - Tiptap escaped footnotes (
[^idor]became\[^idor\]), so they render as literal text, and double-escaped entities: became&nbsp;, which renders as the literal text . - Line endings: every engine except patch-on-save changed the line endings of the CRLF file and flattened the mixed file. remark and Milkdown also stripped the BOM. In a Windows repository that shows up as every line changed in
git diff.
Noise
- Tables re-padded (remark, Milkdown, Tiptap). List markers normalised:
+and-became*in remark, prosemirror-markdown and Milkdown, and-in turndown. The1. 1. 1.list was renumbered1. 2. 3.by all five. - Hard breaks converted: two trailing spaces became a backslash in remark, prosemirror-markdown and Milkdown, and the reverse in turndown and Tiptap.
The naive CodeMirror 6 save is byte-perfect on LF files, but rewrites 100% of the CRLF file and a third of the mixed file, because the editor normalises line breaks to LF. A source editor needs a save layer too. An editor that sets a lineSeparator would keep a uniformly CRLF file; a mixed file still needs per-line handling.
Result B: real edits
Edits issued the way rendered-view widgets would issue them, as minimal source patches, through the patch-on-save reference.
| File | Edits | Result | A naive save would change |
|---|---|---|---|
| 01 Pentest findings (LF) | Tick a task, edit a table cell, insert a table row, fix a word | 3 changed, 1 added | 3 lines |
| 02 AI code audit (CRLF + BOM) | Edit a table cell, insert a 2-line paragraph | 1 changed, 3 added. BOM kept, new lines written as CRLF (115 to 118 CRLF) | 115 lines |
| 03 SOC 2 (mixed CRLF/LF) | Tick a task, edit a cell in an 8-column table | 2 changed. The 24 CRLF and 47 LF lines stay as they were | 25 lines |
| 06 Edge cases (no final newline) | Edit a cell in a pipe-less table, edit the last line | 2 changed. Still no final newline | 2 lines |
Every byte outside the edited lines (BOM, per-line endings, trailing spaces, tabs, final newline state) is copied from the original.
Result C: speed
| Operation (ms, one run each) | 1 MB | 5 MB | 20 MB |
|---|---|---|---|
| lezer full parse (CodeMirror's markdown parser) | 58 | 217 | 873 |
| Patch-on-save: open, 3 edits, save | 9 | 52 | 222 |
| markdown-it render | 65 | 381 | 1,216 |
| marked render | 51 | 287 | 906 |
| prosemirror-markdown parse | 39 | 184 | skipped |
| remark parse | 1,058 | 18,747 | skipped |
| Milkdown load and save | 20,596 | skipped | skipped |
| Tiptap load and save | 111,611 | skipped | skipped |
Node 26 on Apple silicon (M5). Milkdown and Tiptap ran in jsdom, which is slower than a browser, so treat their absolute times as pessimistic; the gap is the signal. Timings are single runs and vary by about a factor of two between runs. The round-trip, construct and edit results are deterministic.
Limits of this benchmark
- Six files, deliberately dense with hard constructs. Plain prose with a few headings round-trips much better in every engine (see file 05, where engines changed 7 to 16% of lines).
- "Lines changed" counts formatting changes too. Re-padding a table is not content loss, but it makes review and
git diffnoisy. - UTF-8 only. UTF-16 and legacy encodings are not tested.
Reproduce it
git clone https://github.com/Katta041/markdown-roundtrip-benchmark.git
cd markdown-roundtrip-benchmark
npm ci
npm run bench -- --quick # a few seconds: round trip, edits, 1 MB timings
npm run bench # full run, about 3 minutes: adds 5 MB and 20 MB timings Requires Node 24 or later. The run writes results/RESULTS.md and the exact bytes each engine saved to results/output/<engine>/, so you can diff it yourself. To add an engine or correct an unfair configuration, open a pull request: an adapter is one object with a roundTrip function. Everything is on GitHub, MIT licensed.
Results as published in the repository on 2026-10-01.