From editor JSON to fast, crawlable pages
This is part 3 of a series on how this blog is built. Part 2 was about the editor. This part follows a post after I press save: where it goes, and how it becomes the page you are reading.
Store the document, not the HTML
The editor could hand me finished HTML, and I could store that. I store the document as JSON instead. A paragraph with one bold word looks like this:
{
"type": "paragraph",
"content": [
{ "type": "text", "text": "Posts are " },
{ "type": "text", "text": "stored", "marks": [{ "type": "bold" }] },
{ "type": "text", "text": " as JSON." }
]
}There are three reasons I prefer this.
It is what the editor reads. Opening a post for editing loads the JSON straight back in. Nothing has to be parsed from HTML and guessed at.
The HTML can change later. If I decide images should render differently, I change the renderer and every post updates. With stored HTML I would have to rewrite old posts.
It is safe by construction. The renderer only outputs nodes and marks it knows. There is no path for arbitrary HTML to reach the page.
The table
The whole document lives in one jsonb column on the posts table:
export const posts = pgTable(
"posts",
{
id: uuid("id").primaryKey().defaultRandom(),
slug: text("slug").notNull().unique(),
title: text("title").notNull(),
excerpt: text("excerpt"),
content: jsonb("content").$type<JSONContent>().notNull(),
coverImage: text("cover_image"),
status: postStatus("status").notNull().default("draft"),
publishedAt: timestamp("published_at", { withTimezone: true }),
readingTime: integer("reading_time").notNull().default(1),
// Edits to a published post that are saved but not live yet.
draft: jsonb("draft").$type<PostDraft>(),
// ...
},
(t) => [index("posts_status_published_at_idx").on(t.status, t.publishedAt)],
);Drizzle's $type gives the column a TypeScript type, so post.content is a Tiptap document everywhere in the code, not unknown. The draft column is for part 4.
One extension list for both sides
Tiptap can turn a document into HTML without a browser. The catch is that it needs to know your schema: which nodes exist and how each one renders. If the server's idea of the schema differs from the editor's, a post looks one way while you write it and another way once published.
So there is exactly one function that defines the schema, and both sides call it:
// Extensions shared by the editor (client) and the HTML renderer (server).
// Client-only behaviour (placeholder, slash menu) is added in the editor.
export function getBaseExtensions() {
return [
StarterKit.configure({
codeBlock: false,
heading: { levels: [2, 3, 4] },
link: {
openOnClick: false,
autolink: true,
HTMLAttributes: { rel: "noopener noreferrer nofollow", target: "_blank" },
},
}),
CodeBlockLowlight.configure({ lowlight: detectingLowlight, defaultLanguage: null }),
CaptionedImage.configure({ HTMLAttributes: { loading: "lazy", decoding: "async" } }),
CaptionedYoutube.configure({
nocookie: true,
controls: true,
HTMLAttributes: { class: "youtube-embed" },
}),
Highlight.configure({ multicolor: true }),
Marker,
TextStyle,
Color,
];
}A few choices are visible here. Headings start at level 2 because the post title is the page's only h1. Images are lazy-loaded. YouTube embeds use the no-cookie domain. Each of these is set once and applies in the editor and on the published page.
The renderer is then very short:
export function renderContent(doc: JSONContent): string {
if (!doc?.content?.length) return "";
const html = generateHTML(doc, getBaseExtensions());
return highlightCodeBlocks(html);
}[Screenshot to add: The same post side by side: in the editor on the left, published on the right]
Highlighting code twice
That second line, highlightCodeBlocks, exists because of a detail that is easy to miss.
In the editor, code is coloured with decorations. They are drawn on top of the text and are not part of the document. So when the document is serialised to HTML, the code block comes out as plain text inside pre and code, with no colours at all.
The fix is to highlight again on the server, with the same library and the same detection rule:
function highlightCodeBlocks(html: string): string {
return html.replace(
/<pre><code(?: class="language-([\w+#-]+)")?>([\s\S]*?)<\/code><\/pre>/g,
(match, lang: string | undefined, body: string) => {
const code = decodeEntities(body);
try {
const tree =
lang && !isUnsetLanguage(lang) && lowlight.registered(lang)
? lowlight.highlight(lang, code)
: detectLanguage(code);
const cls = lang ? ` class="hljs language-${lang}"` : ` class="hljs"`;
return `<pre><code${cls}>${toHtml(tree)}</code></pre>`;
} catch {
return match;
}
},
);
}Two things matter here. The code has to be decoded first, because the serialised HTML has < where the source has <, and a highlighter fed escaped text colours the wrong things. And if highlighting throws for any reason, the block is left as it was: a plain code block is better than a broken page.
The colours themselves are CSS classes in the global stylesheet, so the page ships no highlighting JavaScript to the reader.
[Screenshot to add: A highlighted code block on a published post]
Reading time and excerpts
Both are worked out from the document when a post is saved, and stored, so list pages never have to load the full content.
The reading time is a word count divided by 220 words a minute. Getting the words out of a document has one subtlety:
// Inline nodes are concatenated as they are (a bold word in the middle of a
// sentence is several text nodes); only block boundaries become a space.
export function extractText(node: JSONContent | null | undefined): string {
if (!node) return "";
if (node.type === "text") return node.text ?? "";
if (node.type === "hardBreak") return " ";
const children = node.content ?? [];
const inline = children.some((child) => child.type === "text" || child.type === "hardBreak");
return children.map(extractText).join(inline ? "" : " ");
}The obvious version joins everything with a space. That breaks words: a word with one bold letter is several text nodes, and joining them with spaces splits it into pieces, which then shows up in the excerpt. Joining inline nodes with nothing and blocks with a space gets both right.
If I leave the summary field empty, the excerpt is the first 200 characters of that text, cut at a word boundary.
Lists that do not load the posts
A list of posts needs a title, a date and a summary. It does not need the full document, and the document is by far the largest column. Every list query selects a fixed set of small columns:
// Everything a list needs. Lists never select `content` (the full Tiptap
// document), so their cost does not grow with the length of the posts.
const summaryColumns = {
id: posts.id,
slug: posts.slug,
title: posts.title,
excerpt: posts.excerpt,
coverImage: posts.coverImage,
status: posts.status,
publishedAt: posts.publishedAt,
readingTime: posts.readingTime,
createdAt: posts.createdAt,
updatedAt: posts.updatedAt,
};What search engines get
Because the HTML is produced on the server, a crawler receives the complete article in the first response, with no JavaScript needed to see it. On top of that, each post page provides:
A title and description, with separate search fields I can fill in when the post title is not the best search title.
Open Graph and Twitter tags, so links unfurl properly.
A generated preview image with the post title on it.
BlogPostingstructured data, with the publish and update dates.A canonical URL.
The site also serves a sitemap and an RSS feed, both built from the same query for live posts.
The typography
The published page uses Tailwind's typography plugin, with my own adjustments on top, and the editor's content area uses the same classes:
attributes: {
class: "prose prose-neutral prose-article max-w-none min-h-[50vh] pb-24 focus:outline-none",
},This is the other half of keeping the two sides in sync. The shared extension list makes the HTML the same; the shared classes make it look the same. What I see while writing is, as closely as I could manage, what gets published.
Part 4 picks up at the moment of publishing: drafts, scheduled posts, and how a new post appears on the site without a redeploy.
The whole series
From editor JSON to fast, crawlable pages (this post)
A one-person admin: Google sign-in, a sign-in log and direct uploads