Long-form research
threads.

A series is a research question that turned out to need more than one article. Each thread below states a question, works through it in order, and names the software the argument is implemented in — so a reader can check the claim against a repository rather than take it on the strength of the prose.

Series differ from standalone articles in one respect that matters for reading order: later instalments assume the vocabulary established earlier. A comparison article can be read cold; the second part of a thread usually cannot. Where a series is still open, the question has not been settled and the instalments should be read as working notes rather than conclusions.

Prompt Theory

LLM Cognition & Prompt Theory

How prompts became first-class engineering artefacts. The argument for declarative specification, the lifecycle problem, and the open questions in the field.

The thread starts from an observation that is easy to state and awkward to act on: prompts are load-bearing production artefacts that are handled worse than any other part of the stack. They are interpolated into source at call sites, edited without review, shipped without a version, and compared without a baseline. Every failure mode that version control and type systems were invented to prevent recurs, in an artefact that directly determines system behaviour.

What the articles have settled so far is that the problem decomposes into three separable questions. Specification asks whether a prompt can be written once and executed faithfully across providers. Lifecycle asks how a prompt gets from a codebase into a reviewable, hash-pinned artefact and back again. Optimisation asks whether the wording can be improved automatically against traffic rather than by intuition. They are independent: a team can solve any one of them without the others, and conflating them is why prompt tooling discussions tend to talk past each other.

  1. 01 Formalising Prompts as First-Class Research Objects
  2. 02 Prompt Lifecycle Management: From Extraction to Deployment

Systems Languages for AI

Safe & Verifiable Computing

Why Rust, Zig, and Go are the default for AI infrastructure, and what the trade-offs look like in practice.

AI infrastructure has converged on a small set of systems languages, and the usual explanation — that they are fast — is the least interesting one. The thread argues that the real driver is falsifiability. A runtime with deterministic memory behaviour produces measurements that can be compared across runs; a runtime with a garbage collector produces measurements that have to be argued about. When the claim under test is about tail latency or isolation overhead, the language choice decides whether the claim can be tested at all.

The trade-offs are not free and the articles say so. Rust buys memory safety at the cost of a borrow checker that shapes the architecture, sometimes badly. Zig buys explicit control and comptime metaprogramming at the cost of a small ecosystem and an unstable language surface. Go buys operational simplicity and gives back predictable pause behaviour. The thread treats these as engineering trades with a right answer per problem, not as a ranking.

  1. 01 Why We Write AI Infrastructure in Rust (and Zig, and Go)
Implemented in: zviznumaperfmemistaliath

Threads not yet assembled

Two more questions have enough published material to become series and are not yet ordered into one. The first is agent memory: the distinction between context and memory, the typed-schema argument against treating memory as a single vector bag, and the forgetting problem nobody has a satisfying answer to. The second is the natural-language-to-solver bridge, where the interesting result is negative — language models are good at translating a problem statement into a formal encoding and cannot be trusted to solve the encoding, which argues for composition rather than replacement.

Until they are ordered, the material lives as standalone articles in the blog, grouped by research pillar. The glossary is the fastest way into the vocabulary a series assumes, and Compare places each implementation against the alternatives readers actually evaluate.