How this site is built

This is the third version of dustinedwards.info. The first launched in 2006, the second, on WordPress, in 2016, and this one in 2026.

Everything below is generated from this repository's own configuration and checked against it in both directions on every build. If a binding is added and this page is not regenerated, the build fails. The one thing no generator can produce is why each piece is load-bearing, so those notes are written by hand and reconciled against the bindings they describe.

Runtime

The compatibility date, the flags the Worker runs under, and the Node version the build is pinned to.

Compatibility date
2026-09-01
Compatibility flags
none
Node
24.21.0

Bindings

Every resource this Worker holds a handle to, with what it is for and the measurement that put it there rather than a logo wall.

12 resources.

Retrieval and answer generation for Ask, on built-in storage with uploaded markdown rather than the crawler.

Why it is load-bearing. It sits above classic search rather than replacing it, because the two fail on OPPOSITE inputs. Measured on the same corpus: classic found 2 things AI retrieval missed, all exact tokens, and AI retrieval found 3 that classic missed, all natural-language questions. Neither layer subsumes the other.

ANALYTICS (analytics_engine)

Per-path traffic, one row per HTML response, written server-side from the Worker's response path. Path, referrer host, country and a coarse mobile bit. No cookie, no IP, no client identifier of any kind, so nothing joins two requests together.

Why it is load-bearing. It is here because the alternative was worse. Cloudflare Web Analytics reports the same per-path data for free, but its beacon is a third-party script, and this site's CSP carries strict-dynamic, under which host allowlists are ignored: the script would have to be handed the per-request nonce and connect-src would have to gain a new host. Measured against that, writing the row server-side costs the security boundary nothing and ships no client JavaScript. Workers metrics cannot substitute, because that dataset has no path dimension at all, and the invocation logs that did carry one are now off: they persisted cf-connecting-ip and the session cookie for seven days and no setting redacts a field.

ASSETS (assets)

The static files under public/, served without a Worker invocation.

Why it is load-bearing. The binding has exactly one method, fetch(), so a Worker can serve any path it is given and discover none of them. That is why the media rebuild reads a committed manifest: the one thing it cannot do for itself is find out which static files exist.

BROWSER (browser)

A headless browser that prints the CV to PDF. The Worker calls it after each CV save and keeps one PDF in the derived bucket.

Why it is load-bearing. The CV is edited through Carrel and goes live with no deploy, so a PDF committed to the repository went stale on every save and failed CI until someone re-rendered it by hand. The PDF is derived from the CV rows, so it is rendered where the rows live, stamped with a fingerprint of the data it was drawn from, and a health check compares that fingerprint with the CV the site shows. A failed render never undoes a save: the check goes red and the watcher re-renders.

DB (d1)

SQLite at the edge. Holds every post, the tag graph, the derived search corpus, the media index and the llms.txt row.

Why it is load-bearing. The media index lives here because R2 paginates in key order and promises nothing else. Sort by date, filter unused, search by name, count by type: each is a query, and building one over a key-ordered iterator means listing the whole bucket per request. A capability decision, not a scale one.

ASK_BUDGET (durable_object)

The per-IP burst limit and the site-wide daily ceiling for Ask, on the synchronous SQLite storage API.

Why it is load-bearing. Both cheaper mechanisms were built first and both leaked. The GA ratelimit binding, configured at 5 per 60 seconds and attacked with 12 concurrent requests, refused 1, then 2, then 9, then 0 across four runs. A Durable Object using async storage.get then storage.put allowed 8 through a ceiling of 3, because a read and a write spanning an await is not atomic. The synchronous API allowed exactly 3, which is why the class is registered as new_sqlite_classes.

IMAGES (images)

Thumbnail and content-width transforms, computed on request from the original. Nothing is written back.

Why it is load-bearing. The BINDING and not the /cdn-cgi/image/ URL syntax, and that is forced rather than preferred: the URL interface answers 404 with Cloudflare error 1042 on this hostname, because transformations there require a zone and this site is served from workers.dev.

APP_KV (kv)

Better Auth sessions, and the Ask answer cache keyed by a SHA-256 of the normalized question.

Why it is load-bearing. The cache sits in front of the daily ceiling rather than behind it, so a repeated question costs nothing and consumes no budget. The ordering is the design: a hit reaches no model.

dustinedwards-media-events (queue_consumer)

R2 event notifications from both buckets, consumed to derive the media index rows.

Why it is load-bearing. The write path is a notification into a queue, never a dual write. The Worker writes only to R2 and the consumer re-derives the row from the object as it is now, never branching on what the message claimed, which is what makes replay and out-of-order delivery converge.

MEDIA (r2)

Uploaded originals, content-addressed as sha256(bytes) truncated plus the extension.

Why it is load-bearing. Content addressing was decided on security, not tidiness. The old key was derivable from public information, so a draft's images were reachable by guessing a slug the sitemap publishes plus eight hex characters.

MEDIA_BACKUP (r2)

A mirror of MEDIA. Every uploaded original has a byte-identical twin here, written by copy and never by the site itself.

Why it is load-bearing. The realistic way these bytes are lost is this site deleting them: the OG prune, the media delete action and the R2-wins reconciliation can each remove an object with no undo. So no code path deletes from this bucket, check:destructive sweeps for one, and health compares every object against its twin. It replaced an acceptance that held only while MEDIA was empty, which expired on the first upload.

OG (r2)

Social cards, one per post, rebuilt by command, and the CV's PDF, one object replaced each time the CV is saved.

Why it is load-bearing. A second bucket exists because the two split on LIFECYCLE rather than on what the UI calls them. MEDIA is irreplaceable; OG is regenerable, so it is safe to empty and is deliberately outside the backup path.

Schema

Hand-written migrations, applied in order. drizzle-kit is deliberately not a dependency, and because the database export command is broken on this schema, this directory is the only copy of the table definitions that exists anywhere.

25 migrations.

  • 0001_init.sql
  • 0002_blog_content.sql
  • 0003_post_toc.sql
  • 0004_blog_discovery.sql
  • 0005_og_image.sql
  • 0006_search_index.sql
  • 0007_media.sql
  • 0008_llms_seed.sql
  • 0009_media_index.sql
  • 0010_media_role.sql
  • 0011_media_trash_tags.sql
  • 0012_drop_posts_category.sql
  • 0013_content_hashes.sql
  • 0014_webmentions.sql
  • 0015_zero_result_queries.sql
  • 0016_post_head_blocks.sql
  • 0017_post_changelog.sql
  • 0018_post_backlinks.sql
  • 0019_procedures.sql
  • 0020_publications.sql
  • 0021_pages.sql
  • 0022_cv.sql
  • 0023_dictionary.sql
  • 0024_roster.sql
  • 0025_phages.sql

Gates

Checks that run before anything ships. The list is derived from the scripts themselves rather than maintained beside them, so a gate that is added and forgotten is not possible.

22 checks.

  • check:ask-guards
  • check:backup
  • check:browser
  • check:content
  • check:contrast
  • check:destructive
  • check:diagrams
  • check:enhance-a11y
  • check:features
  • check:fonts
  • check:headers
  • check:links
  • check:machine-readable
  • check:migrations
  • check:policy
  • check:protocols
  • check:restore
  • check:secrets
  • check:tests
  • check:types
  • check:urls
  • check:worker

Dependencies

The runtime dependencies. Build tooling is excluded: this is what serves the site, not what assembles it.

35 runtime dependencies.

  • @codemirror/commands 6.10.4
  • @codemirror/lang-markdown 6.5.1
  • @codemirror/language 6.12.4
  • @codemirror/state 6.7.1
  • @codemirror/view 6.43.7
  • @dustinedwards/site-api github:DrDustinEdwards/site-api#v0.1.0
  • @lezer/highlight 1.2.3
  • @observablehq/plot 0.6.17
  • @shikijs/rehype 4.3.1
  • @shikijs/transformers 4.3.1
  • better-auth 1.7.6
  • drizzle-orm 0.45.2
  • enarratio 0.1.0-alpha.8
  • github-slugger 2.0.0
  • gray-matter 4.0.3
  • hast-util-to-string 3.0.1
  • image-size 2.0.3
  • katex 0.16.47
  • linkedom 0.18.13
  • react 19.2.7
  • react-dom 19.2.7
  • react-router 8.3.1
  • rehype-autolink-headings 7.1.0
  • rehype-katex 7.0.1
  • rehype-slug 6.0.0
  • rehype-stringify 10.0.1
  • remark-directive 4.0.0
  • remark-gfm 4.0.1
  • remark-math 6.0.0
  • remark-parse 11.0.0
  • remark-rehype 11.1.2
  • shiki 4.3.1
  • unified 11.0.5
  • unist-util-visit 5.1.0
  • zod 4.4.3

What it does

Nothing in this section can be generated from configuration, so each entry carries an anchor and a gate verifies that the thing the claim is about still exists. It does not verify that the sentence is true, which is the honest boundary of the technique.

35 entries.

Content pipeline

Markdown is the source of truth, D1 owns every read

Posts are markdown files in the repository, and git holds the markdown and nothing else. D1 holds the only rendered copy, and every public read queries D1. Each row records the git blob sha of the source it came from and a hash of its own render, so a ship compares both against a fresh build before writing and names by slug every row that has drifted. A row whose source is unchanged but whose render differs fails the run after the deploy stands, because that is a defect in the pipeline rather than staleness.

One renderer, two writers

The build scripts and the Worker import the same markdown pipeline module. There is deliberately no second renderer, because both writers render the same source independently, the Worker on every save and the build on every ship, and a difference between the two is invisible until something compares them.

  • route /writing/:slug
  • gate check:content

Every post has a markdown twin

The same URL with .md appended returns the source the page was rendered from. Requesting the post with an Accept: text/markdown header returns the same bytes.

  • route /writing/:slug.md
  • gate check:machine-readable

Charts

Charts are content, not a widget

A chart is authored as a container directive over a fenced block of CSV and rendered to SVG as the post is rendered, so it is part of the stored HTML and a reader with JavaScript disabled sees the same chart as everyone else. Colors are CSS custom properties that resolve per theme, which is why one stored SVG serves light and dark with nothing to flash.

  • gate check:content

Diagrams

Diagrams render ahead of time, one drawing per theme

A diagram needs real font metrics, so it cannot render in a Worker. One SVG per theme is produced ahead of time and the stylesheet shows one of them. The rendered post carries the key and the source rather than the drawing, so a font-rendering difference between two machines can never turn into a failed build. Alt text is mandatory.

  • gate check:diagrams
  • assertion in check:diagrams: a diagram with no alt fails

Two FTS5 indexes over one derived table, fused by rank

An FTS5 tokenizer is set per table, not per column, so one index cannot serve both exact-token lookup and prose relevance. There are two: unstemmed for titles and tags, Porter-stemmed for bodies. The two ranked lists are merged by reciprocal rank fusion, never by raw score, because scores from two tokenizers over two average document lengths are not comparable on value.

Results are section-grained

One post yields a document record plus one per heading, so a hit deep-links to the heading that answers it rather than to the top of the page.

A filter with no text takes a different query path

A bare year or a tag chip has nothing to give FTS5, so those queries run the filter SQL straight over the derived table, date ordered, document records only. Returning every section of one post for a bare year would present the corpus at several times its real size. The snippet carries no highlight, because nothing was matched and marking anything would claim a match that never happened.

The same URL returns JSON under content negotiation

Requesting the search page with Accept: application/json returns the same query against the same index, with Vary: Accept. It is not a second API. Negotiation runs in middleware, because a document route's loader cannot return a raw Response.

Ask

A streamed AI answer with citations that deep-link

Retrieval and generation sit strictly above classic search and never block it. Citations map back to heading anchors, and the answer is written with textContent rather than as HTML, because it is generated text and the one thing known about it is that we did not write it.

  • route /search/ask

The two search layers fail on opposite inputs

Measured on the same corpus: classic keyword search found two things AI retrieval missed, both single tokens naming a heading almost verbatim, and AI retrieval found three that classic missed, all long natural-language questions. Neither layer subsumes the other, which is the case for keeping classic first and Ask on top rather than replacing one with the other.

Three cost gates in front of the only metered endpoint

A per-IP burst limit, then an answer cache keyed by a hash of the normalized question, then a site-wide daily ceiling. The ordering is load-bearing: the ceiling sits after the cache, so a cache hit costs nothing and consumes no budget. A refusal is a 429 with Retry-After and reaches no model.

  • route /search/ask

Drafts can never reach the AI index

The classic index filters drafts at query time. The AI index cannot, because it has no per-item status a query can filter on, so exclusion happens at upload time. A draft uploads nothing and actively removes anything the post already had, because skipping the upload alone would leak on the unpublish path: a post published, indexed, then withdrawn would stay answerable forever.

  • gate check:ask-guards
  • assertion in check:ask-guards: a withdrawn post is not Ask-publishable

Rate limiting

Durable Objects, because the platform limiter does not count

Both cheaper mechanisms were built first and both leaked. Cloudflare's GA ratelimit binding, configured at five requests per sixty seconds and attacked with twelve concurrent requests, refused one, then two, then nine, then zero across four runs. It is documented as permissive and eventually consistent: it sheds sustained load rather than counting. A Durable Object using the asynchronous storage API allowed eight through a ceiling of three, because a read and a write spanning an await is not atomic even inside a single-threaded object. The synchronous SQLite API landed on its ceiling exactly, which is the whole reason the class is registered as a SQLite-backed one.

  • route /search/ask

Admin editor

The markdown is committed before the database is touched

A save commits exactly the markdown, through the Git Data API, and writes D1 only after that commit lands. Files are the source of truth, so nothing reaches the database that is not already in git, and the ordering is the whole guarantee: a database row can always be rebuilt from the repository, and a commit can never be rebuilt from a row. If the repository host is unreachable the save fails whole; there is no database-only write and no reconcile-later queue.

  • gate check:policy

The editor records the head commit and refuses a conflict

The editor notes the head commit when it loads and sends it back on save. If the branch moved, the save is refused with the divergence named and no commit is created. The reference update is never forced, so the editor cannot overwrite a change it did not see.

  • gate check:worker

Version history restores by writing a new commit

History lists the commits touching a post's file, shows a diff per commit, and restores by reading the file at an old commit and putting it back through the ordinary save path. A restore is therefore a new commit, never a force push and never a rewrite, and the gates run again on restored content so an old commit cannot bypass a newer rule.

  • gate check:content

One feedback slot, four named outcomes

A first publication used to complete in silence, which is a design failure for the one act the system reserves to a human. The slot names the transition rather than the action: saved, first publication, republished, unpublished. The transition is decided by the publish policy module and not by the interface, because current state cannot tell a first publication from a republication: a post marked draft is either brand new or previously withdrawn.

  • gate check:policy

Operator API

Agents write through the same path a human does

A bearer-token JSON endpoint exposes the editor's save machinery to non-browser callers. It calls the same save and delete functions the browser action calls, so an agent gets the same schema gates, the same style check, the same single atomic commit, the same database sync and the same index sync. Nothing in the operator layer touches the repository, the database or the AI index directly.

  • route /api/operator
  • gate check:policy

An agent may not perform a post's first publication

An operator may create, edit, unpublish and republish. It may not make a post public for the first time; that is refused with a 403 naming the policy. The durable fact lives in frontmatter, is stamped once by the save path, and is read only from the committed file and always overwritten on the way out, because otherwise an agent could publish any draft by asserting the very fact the gate checks.

  • route /api/operator
  • assertion in check:policy: readState reads a previously published post

Media

R2 is the store, D1 is a rebuildable index

R2 holds the bytes and wins every conflict: a row with no object is deleted, an object with no row is backfilled, never the reverse. The index exists in D1 because R2 paginates in key order and promises nothing else, so sorting by date, filtering unused, searching by name and counting by type are each a query, and building those over a key-ordered iterator means listing the whole bucket per request. A capability decision, not a scale one.

  • route /media/*

Keys are content-addressed and carry their dimensions

A key is a truncated hash of the bytes plus the extension, which was decided on security rather than tidiness: the old scheme was derivable from a slug the sitemap publishes plus eight hex characters, so a draft's images were guessable. The key also carries the pixel dimensions, so the build and the Worker resolve image sizes from the string rather than from two different stores that could disagree.

  • route /media/*
  • assertion in check:tests: the fixture produced both resolutions and refusals

Transforms are computed on request and never written back

One object per image, every size derived from it, which keeps the bucket free of variant sprawl and is why it needs no lifecycle rules. It uses the Images binding and not the URL transform syntax, and that is forced rather than preferred: the URL interface answers 404 with error 1042 on this hostname, because transformations there require a zone and the site is served from a workers.dev address. Thumbnails serve one format unconditionally, after a negotiated version turned out to poison a cache whose key carried no headers.

  • route /media/*

Theming

The theme is resolved server-side, so nothing flashes

The choice lives in a cookie, is read in the root loader, and the correct attribute is written into the first byte of HTML. There is no inline script and nothing is corrected after paint. System preference is the absence of the attribute rather than a third value.

  • route /theme
  • gate check:contrast

The toggle works with scripting off

It is a real form posting to a route that sets the cookie and sends the reader back. The client enhancement only removes the round trip. It is deliberately not a radio group, because a real radiogroup owes arrow-key roving focus that cannot be delivered without script.

Every color pair is recomputed from the shipped stylesheet

The gate reads the token values back out of the stylesheet and recomputes the whole contrast matrix, so a tuned hex moves one side of the comparison and fails. The pairs and thresholds are transcribed as token names rather than values, so nothing in the gate restates a color.

  • gate check:contrast
  • assertion in check:contrast: dark (prefers-color-scheme)

Brand mark

The mark is inline SVG

One component collapses the four SVG files into a path list and a viewBox, because the files differ in exactly two ways. The purple paths carry no fill and take a brand custom property instead, which is the entire dark-mode story for the mark.

  • route /

Machine surfaces

Feeds, a sitemap, and a map of the site for agents

RSS and JSON feeds, a sitemap, robots, and a plain-text map of the site with a full-text companion carrying every published post in one document. ONE FEED POLICY: both feeds read the same query with the same visibility predicate and the same cap, and carry the whole post rather than a summary, differing only in representation because an RSS reader renders markup and a JSON consumer is usually a program. The map is served from a database row rather than from the file in the repository, and the inlined fallback is imported from that same file so the two cannot drift.

Search is exposed over the Model Context Protocol

The same corpus is reachable as an MCP tool, so an agent can search the site without scraping it. It needs no authentication and returns the same section-grained keys the rest of the site uses. It shows the same retrieval characteristic as the AI answer layer: it finds passages that answer a question and misses short exact tokens the keyword index finds immediately, which is stated rather than hidden.

Backups

Per-table exports, because the platform export is broken here

The database export command fails outright on a schema containing FTS5 virtual tables and writes nothing, so backups are taken per table with the schema excluded, never touching the virtual tables or their shadow tables. The table list is derived from the migrations directory, and the check fails in both directions. It found a real omission on its first run.

  • gate check:backup
  • assertion in check:backup: sqlite_master returned no tables

Counting an FTS5 index cannot detect that it is empty

A count on an external-content FTS5 table reads through to the content table, so it reports rows that the index itself no longer holds. Measured twice: with the index emptied, the count still read seven while the size shadow table read zero and matching returned nothing. An integrity check also passed on the broken index. The shadow table is what gets counted, and the repair for a corrupt index is a rebuild, never a delete.

  • gate check:backup

Safety rails

URL protocols are allowlisted in the schema as well as the renderer

The render-layer guard only walks nodes the body renderer produced, so frontmatter bypassed it entirely: a curated link field took a generic URL validator, which accepts javascript:, and rendered as a live public anchor. Both that field and the cover path now call the same predicate the renderer uses rather than reimplementing it. That is asserted against the source, not only through behavior, because the two come apart: the cover path spent a while refusing bad protocols as a side effect of a stricter rule about paths, which every behavioral case passed and no gate could see.

  • gate check:urls
  • assertion in check:urls: frontmatter blocked cases were exercised
  • assertion in check:urls: schema: ${label} calls isAllowedUrl, not a reimplementation

Rules stated twice are checked against each other

Some rules cannot be collapsed into one function because they are expressed in two languages: the public visibility predicate exists once as a query-builder condition and once as a hand-written SQL string over a different table. Those are run against a fixture of post states and must admit the same rows. The same suite proves the column schema agrees between the schema module, the migrations and the deployed database.

  • gate check:tests
  • assertion in check:tests: both predicates admit the same posts

This page

The stack half is generated

Every binding, migration, gate and dependency above is emitted from the repository's own configuration when the site is built. Nothing in the generator or on this page restates a name.

The feature half is anchored, because it cannot be generated

Sentences like this one are in no configuration file, so each carries an anchor: a route, a gate, or a specific assertion inside a gate. The check verifies the anchored thing still exists. It cannot verify that the sentence is true, which is stated in its header rather than implied, and a claim whose only anchor would be an external decision record is refused outright, because that would be an unverifiable claim wearing the costume of a verified one.

A tradeoff in the security headers

The content security policy is enforced, not merely reported. One part of it is a compromise rather than a clean win, and the compromise is worth stating plainly.

Every response carries a one-time number that scripts on the page must quote to be allowed to run. That number is generated per response.

Seven pages of this site are cached at Cloudflare's edge and served to everyone from the same stored copy for up to ten minutes. The number is part of that copy, so visitors served from one cache entry share it.

We took that trade deliberately. The alternative is to stop caching those pages, which would make every reader wait for the origin on every visit.

It is acceptable only because these pages carry no writing from anyone but me. There are no comments and no user submissions, so there is nowhere for a stranger's script to get in and use the shared number.

AI disclosure

This site is built with AI assistance and says so here rather than leaving you to guess, because a site about how it is built owes you that before it owes you anything else.

The code and the prose here are written with AI assistance. The model is Anthropic's Claude, driven through Claude Code, and it reaches this site through exactly the same publishing API and the same gates a person does.

Nothing goes public on a model's say-so. Making a post public for the first time is reserved to me and the reservation is enforced in code, not asked for in a prompt: an agent that attempts it is refused by name.

Every published post has been read and edited by me before it went live, so what you are reading is human-reviewed writing I am answerable for rather than model output passed straight through.

Where a provider marks its model's output in a machine-readable way, that mark does not survive being edited, so it is not something you can check on this page. The review above is the guarantee instead, which is why it is stated as a practice and not as a badge.

How the site is run

Three other systems of mine sit behind or beside this one. None is part of this repository.

Capsid is the site's memory and job queue. It stores the project's decisions and current state, and it hands work from a conversation to a working session as a signed job that ends in a pull request.

Carrel is my private writing hub, behind a login, where articles are written before they are published here. A finished article reaches this site through a small keyed API.

Capsid's source is public on GitHub.

Capsomer is the shared design system this site will adopt.

What was not adopted

Anyone can list what they shipped. A refusal is a decision that was made and recorded: not a thing nobody got to, but a thing somebody ruled out, with the reason beside it.

5 entries.

Vectorize (Refused)
Rank fusion over two FTS5 indexes answers this corpus in about fifteen lines and no vectors. Zero-result queries are the cheapest content-gap signal the site has, and the only evidence that would justify reopening the question.
Images storage (Refused)
Transforms are used; storage is not. It bills in $5 increments per 100,000 stored images, for a library of dozens.
Soft delete and a trash prefix (Refused)
A team concern. The audience is one person, and R2 wins every conflict already, so a deleted row that should not have gone is restored by the next rebuild.
Media versioning (Refused)
Keys are content-addressed, so a changed image is a different key by construction. Versioning would be a second mechanism for something the addressing already does.
Single-flight on Ask (Refused)
Declined with the bound written into the route. Deduplicating identical questions does not lower the worst case: N addresses spend the same N units asking N DIFFERENT questions. The 200 per day ceiling is the real bound and it is exact.

For what the site records about a visit, and where each of those facts lives in the code, see privacy.