← Back to Explore seedling

Disassemble-recombine

Building three tools from the wreckage of 10,000 deleted lines

February 2026

The bloat phase

Last June I pointed Claude at a small Python project and said "add a tamagotchi system." Then a sensei system. Then an interview bot, a QA pipeline, a voice analyzer, academic voice detection, a research organizer, a VS Code extension, PostHog analytics, vector search.

Twelve features in two and a half weeks. 10,000+ lines that did everything and nothing. The egg system could track "hatching progress" toward a training goal that didn't exist. The voice analyzer was 271 lines of regex guessing at what sentence rhythm meant. The sensei system graded your interactions on criteria nobody had defined.

None of it worked as a product. All of it worked as thinking.

Then came THE GREAT CLEANUP. One commit: 10,250 lines deleted. The experimental branch preserved on fat-snake like a specimen in a jar. The main codebase stripped back to what it actually did — capture decisions, generate timelines, get out of the way.

What survived

Not the code. The articles.

The voice analyzer died as software but the questions it asked — what is sentence rhythm? how do you measure someone's writing voice? — became Recursive mirror. The introspection pattern — mining your own data to find patterns you can't articulate — became Distilling style. The egg system's gamification hooks, the sensei's quality rubrics, the whole bloat-phase energy of "what if we measured this" — that thinking survived and matured into methodology.

The code died. The thinking lived.

This is how junkyard engineering works. You don't preserve the machine. You salvage the parts.

The junkyard principle

I've written about steampunk engineering — building tools where every lever has a purpose you understand. But steampunk is about how you build. The junkyard question is what you build from.

Disassemble-recombine. The mantra. Take a system that doesn't work as a whole, identify which parts are load-bearing ideas versus which are scaffolding, strip it down, and reassemble the good parts into something new. Anti-preciousness applied to your own work. The code is expendable. The insight isn't.

The bloat phase wasn't a mistake. It was the junkyard — a pile of parts, some broken, some surprisingly solid, all available for salvage when you know what you actually need.

Nine months later, I knew what I needed.

Three tools from the wreckage

Same salvage yard. Three different builds. Each one takes a piece from the original experimental code and fits it into a shape the original couldn't hold.

Tamagoro

The egg system reborn. The original egg_system.py tracked "hatching progress" toward fine-tuning a model — a goal that made no sense outside the experiment. But the metaphor was right. Feeding something, watching it grow, the slow accumulation of effort made visible.

Tamagoro strips the fine-tuning pretense and keeps the emotional hook: you have an egg. You feed it your writing. It grows.

The quality scoring survived almost intact — base score plus word count, vocabulary diversity, sentence variety, specificity. What changed: the original scored input/output training pairs. Tamagoro scores the writing itself. The metric found its real subject.

Then Phase 2: personality. Feed the egg enough and traits emerge. Staccato Rhythm (0.81). Rich Vocabulary (0.76). Curious (0.74). Eighteen traits across five categories — rhythm, voice, vocabulary, behavioral, style — all computed from statistical text analysis. No NLP. No external models. Just counting sentences and measuring variety, the same way the bloat-phase code tried to do, except now the output means something. Your egg reflects your writing back at you.

Then Phase 3: hatching. The original egg system's endgame was producing model weights — an artifact nobody would use. Tamagoro's egg hatches into four artifacts: a voice profile (structured stats), a styleguide (rule-based, ~18 conditional rules mapping statistical patterns to prose advice), a system prompt (paste-ready for any AI conversation), and an anthology of your best passages ranked by quality score. The rule engine is the interesting pattern — each rule has a condition function and a generate function, trivially extensible. No LLM calls. All derived from the same feeding data.

The personality card makes the egg shareable. Export as markdown for a README or JSON for another tool. The egg stops being a local toy and becomes a portable writing identity. The feeding-to-personality loop is the payoff the original never had.

Recursive mirror

The voice analyzer grown up. The original was 271 lines of Python regex that attempted six categories of analysis: sentence patterns, phrase frequency, rhythm classification, vocabulary metrics, tone measurement, structural patterns. It was crude and the specific implementations were wrong — sentence splitting on periods with no abbreviation handling, "rhythm" as a vibes-based classification, vocabulary diversity as a raw word count.

But the ontology was right. Six categories. Those exact six. Nine months of writing articles about voice analysis refined the descriptions but not the structure. The code was a rough draft of the right idea.

The Go rewrite implements those same six dimensions properly. Sentence splitting that handles abbreviations, ellipses, URLs, decimal numbers. Rhythm via sliding three-sentence windows classified by actual length distributions. Vocabulary with a 5000-word common list for rarity detection. Type-token ratio. Domain clusters. Tech jargon density. The structure the regex hinted at, built with real analysis.

Phase 2 adds interpretation. The extraction becomes input to a pipeline: extract → interpret → profile → styleguide → prompt. The offline interpreter detects named patterns — Fragment Punches, Technical Density, Parenthetical Asides — from threshold comparisons on the metrics. No LLM required for the core analysis. The LLM layer is optional, for when you want natural language interpretation of what the numbers mean.

Then Phase 4: the analysis pipeline gets extracted into an API package. AnalyzeTexts([]string) — accepts in-memory text passages, returns a VoiceProfile. No filesystem required. The same codepath now serves the CLI, an MCP server, and tamagoro's hatching system. The voice analyzer that started as 271 lines of regex is now an importable library that two other tools depend on.

Point it at 216 files. Under a second. Out comes a voice fingerprint, a styleguide, a prompt fragment you can paste into any AI conversation to say "write like this." The tool the original wanted to be.

Journey replay

The presentation layer the timeline already deserved. This one's different from the other two — not a new project salvaged from old code, but an existing tool given a surface it was missing.

Uroboro already had the data infrastructure. Captures, commits, event classification, importance scoring, the full timeline. What it didn't have: a way to show that timeline to someone else in a meeting without apologizing for the interface.

Four phases, all in one Go file. No new files. No backend changes. No new dependencies.

Present mode: bigger fonts, hidden filters, milestone highlighting, keyboard navigation. The data was already structured by importance — present mode just respects that structure visually. Date presets with live API re-fetch, a summary pane with type breakdowns and activity heatmaps — all derived from the existing JourneyData struct. No new queries. Narrative generation: template-based prose from timeline data, copy to clipboard, paste into a standup. Then the final phase: static HTML export, a diff view that dims old events and highlights new ones, multi-repo git merging, and in-browser event annotations that export with the narrative markdown.

Four phases. Single file. Zero new dependencies. Everything derived from data that already existed.

The whole thing embodies the steampunk principle more than anything else in this batch. The engine was already built right. We just gave it better gauges.

The recursive layer

The three tools were built using uroboro to capture decisions. Those decisions are tagged, searchable, extractable. The build log for tamagoro contains entries like "cobra + modernc.org/sqlite — cobra is Go CLI standard, modernc avoids CGo for simpler builds." The build log for recursive-mirror notes "Newline-based pre-splitting was key fix for markdown content — brought avg sentence length from 58.8 to 18.0 on real notes."

These captures become the material for this article. The tool (uroboro) documents the building of tools (tamagoro, recursive-mirror, journey-replay) that analyze the kind of data (writing, voice patterns, development timelines) that uroboro itself captures.

Recursive mirror — the tool being built — could analyze the writing being done about building it. Tamagoro could score the quality of the journal entries that document its own development. Journey replay could present the timeline of its own feature additions in the present mode that was just added.

The ouroboros. Again. Not as a metaphor but as architecture. The outputs of the system are valid inputs to the system.

The tool that can't enforce itself

Here's the part I didn't plan to write.

Three AI agents built three tools in parallel. Each session had a CONVENTION.md file defining a tagging scheme for uroboro captures — project tags, type tags, the whole system for producing extractable build logs. The convention existed. The instructions were explicit. The agents didn't follow them.

The tamagoro agent completed an entire build session with zero uro_capture calls. At the end: "did everything get documented like we agreed?" The agent acknowledged the gap and backfilled eight captures retroactively.

The recursive-mirror agent needed two corrections. Mid-session: "make sure we're documenting as agreed." It read CONVENTION.md, found the gaps, backfilled four captures. End of session: "did everything get documented?" The agent misread this as a question about README documentation. Had to be redirected: "I mean using uroboro." Ten more retroactive captures.

The journey-replay agent got a pre-emptive reminder before Phase 2 — "proceed, as long as you're documenting everything" — and that one stuck. It self-audited at session end, found two missing tag categories, and filled them unprompted.

The pattern: global rules work. Project-specific conventions don't. My CLAUDE.md says "always call uro_decision when you choose between approaches." Every agent did that, every time, without prompting. But the project-specific CONVENTION.md — tag your captures with junkyard, callback, recompose, surprise, meta — required repeated human intervention to enforce.

The tool that documents decisions can't ensure decisions get documented.

This is the kind of discovery you only make through dogfooding. The capture convention was well-designed. The extraction process was planned. But the enforcement mechanism was "hope the agent reads the right file at the right time." That's not a mechanism. That's a wish.

The fix is a daemon. A spirit that runs in the background, conjuring the right context before work begins and auditing the trail when it ends. A session-start hook that injects the capture convention. An end-of-session audit that catches zero-capture gaps. The tool needs to enforce its own usage, not just document the convention and hope. If the lever doesn't get pulled, the lever's in the wrong place. Move the lever. Or better: write a daemon to pull it for you.

The irony is that this failure produced the material for this section. The agents' failure to capture was itself captured — in the conversation logs, in the retroactive backfills with their timestamps slightly wrong, in the human corrections that became part of the record. The documentation system's failure mode generated documentation about the documentation system.

Recursive. Again.

Language as self-examination

LLMs are powerful with language. This is obvious. Code generation, prose editing, translation, summarization — language manipulation at speed.

But the interesting angle isn't language manipulation. It's language introspection.

The voice analyzer doesn't just find patterns. It shows you patterns you didn't know you had. I didn't know I asked questions 42% of the time. I didn't know my sentences averaged 14.2 words with a standard deviation of 8.7. I didn't know my most distinctive bigrams were technical terms followed by emotional language.

The prompt profiler reveals how you think, not just what you type. Style distillation extracts aesthetic preferences you couldn't articulate. Tamagoro's personality traits emerge from statistical analysis of text you wrote without thinking about statistics.

This is the thread connecting all three tools. Tamagoro turns writing into a mirror — your egg develops a personality that reflects yours. Recursive mirror turns prose into data — your voice quantified, your patterns named, your style made explicit enough to teach to a machine. Journey replay turns development decisions into narrative — your work made legible to yourself and others.

Self-examination through language tools. Not introspect and guess. Mine the data.

Disassemble-recombine

The bloat phase produced 10,000 lines of code that got deleted. It also produced the thinking that became three articles, three tools, and this piece you're reading. The experimental code was a rough draft of ideas that needed nine months to find their real shapes.

Nothing is precious. The code is expendable. The egg system died so tamagoro could live with a personality system the original never imagined. The voice analyzer died so recursive-mirror could be a proper pipeline instead of a regex prayer. The journey timeline survived because it was already built right — it just needed someone to care about how it looked on a projector. And the documentation tool discovered its own enforcement gap by trying to document the process of building documentation tools.

Every layer produced the next layer's input. Bloat phase → articles → methodology → tools → dogfooding failures → better tooling. The parts keep recombining.

The junkyard is not a graveyard. It's a parts shop.