The question
The parent article makes a claim about redistribution — that AI shifted cognitive labor from syntax and convention toward intent and design, and that some brains benefit from this shift more than others. That's a nicer story than "I press enter a lot." So here are the receipts.
I have nine months of git history across personal and professional repos. I have 1,649 uroboro captures logging decisions, blockers, and open questions. I have a satirical tool called keygrave that parses my Claude Code session data and spits out the exact kind of metrics that make conclusions easy.
What does a redistribution look like in the data? What does a rationalization?
The commits
948 commits since May 2025. Personal projects and professional client work combined, filtered to my authorship. I stripped the Co-Authored-By Claude tags from most of them — not because I was hiding anything, but because people were rejecting working code the moment they saw the attribution. The code either works or it doesn't.
The personal project curve looks like a Gartner hype cycle. 205 commits in June — seven tools getting built simultaneously. Uroboro, wherewasi, examinator, qoins, qombow, doc-search. Everything at once. Then the cliff: 18 in July, zero in August, 4 in September. Peak of inflated expectations, trough of disillusionment, painted in commit counts.
Pull the camera back. While personal projects went quiet, client work ramped: 48 commits in September, 205 in October across four simultaneous engagements. The energy shifted domain. Bechtle. Kepler. Foodflou. The combined monthly picture: 63, 205, 18, 4, 52, 205, 34, 14, 178, 175. Not sustained. Not collapsed. Cyclical.
Everyone goes through a hype cycle with new technology. The question isn't whether I had one. It's what survived the trough. The tools that made it out — uroboro, wherewasi — got extracted, refined, shipped as standalone projects. The ones that didn't got left behind. That's just adoption.
The intent layer
Here's where it gets more interesting than commit counts.
Uroboro has an MCP server. Most of the captures — pretty much from the sjiek era onward — were logged by Claude Code, not typed by hand. I watched agentic LLMs use tools, bash scripts, CLI, and a lightbulb went on. CLI-first was already powerful for text manipulation. With a large language model as the operator, it became something else entirely. So uroboro was designed to be frictionless for both humans and AI. Either one can capture a decision, a blocker, an open question, with the same interface.
There's a layer of irony here. The tool that documents my intent is itself largely operated by the AI. But what it captures isn't implementation — it's the forks in the path. What I chose, what I rejected, why. "Zero external dependencies over convenience libraries — control and auditability." "SQLite over Postgres — single file, no daemon, portable." "CLI-first over web UI — composable, scriptable, no browser dependency." 384 documented decisions across 54 projects.
Git tells you what happened. Uroboro tells you what was considered and discarded. A diff shows you the code changed. An uroboro capture shows you the three approaches that were discussed before the code was written, and which one won, and why. During any conversation, I can pull up uroboro context — not a code diff. Intent, not implementation.
The patterns recur across unrelated work. Zero-dependency preference in Go, in bash, in site architecture. SQLite as the coordination layer everywhere. CLI-first interfaces even when a web UI would be easier to demo. These aren't AI defaults. Claude would happily reach for Express or React. These are human preferences the AI learned to follow — and that uroboro made legible.
The structured tagging started in January 2026 — but not because I got disciplined. I tried that. The pre-January captures have freeform tags: philosophy,design,empathy,strategy,metaphors,ux,psychology. Stream-of-consciousness. Useful for archaeology, not analysis.
What changed on January 13 was delegation. Uroboro got an MCP server. Claude Code started calling uro_decision and uro_blocker natively. And when the AI does the capturing, it naturally produces consistent taxonomy — decision, feature, bugfix, blocker. The structured tagging emerged as a side effect of MCP integration. The discipline came from the tool, not from me.
That's the redistribution in miniature. I was bad at the bookkeeping. The AI is good at the bookkeeping. Delegate. Now the intent layer is well-instrumented not because I became more organized, but because I stopped pretending I would be. The human handles the selection — which fork, which trade-off, which pattern. The AI handles making that selection legible. Both roles play to strength.
And then the tool grows from its own use. While writing this essay, I used uroboro to investigate uroboro's own history — tracing when the structured tagging started, comparing pre-January captures to post-January ones. To do that, I had to drop into raw SQLite queries because the MCP server didn't expose date-range filtering or tag taxonomy analysis. The tool I built to capture intent couldn't answer questions about its own intent layer.
So the gaps became the next feature. since/until filtering. Project-scoped queries. Tag analysis. The implementation is syntax and convention — Claude writes that. But the recognition that uro_search needs date-range support? That came from using the tool in anger and falling into the hole. You can't spec that in advance. You don't know the gap exists until you're in the middle of the question it can't answer.
This is what the redistribution looks like at full recursion. Build a tool with AI. Use the tool to examine your own process. Hit the tool's limits. Recognize the limits because you understand what you were trying to ask. Have the AI implement the fix. Use the improved tool to ask the next question. The human contribution at every step is the same: knowing what to reach for. The AI contribution at every step is also the same: building the thing you reached for. Neither part works without the other. The meta-meta-meta-commentary is becoming a core identity, and I've stopped fighting that.
The sphinx
During the June peak — while all those tools were getting built with AI running full tilt — I also built sspphhiinnxx.
sspphhiinnxx is a methodology. For every project built with AI, the sphinx generates proof-of-understanding riddles specific to that project. Build qombow — an image compositing tool — and the sphinx asks eleven riddles. "How does each pixel in your composite receive its final value?" "Rebuild the mask expansion logic from scratch." "Can you rebuild qombow from nothing but understanding, without referencing a single line of your original code?"
The tagline: "Ship fast by day. Understand deep by night."
I wrote this during the peak. The same person pressing enter thousands of times in Claude Code was building a system that demands you prove you understand what you just shipped. At the time, that felt like responsible engineering. The intellectual conscience. Learn or die.
Nine months later, I think the sphinx was asking the wrong question — or at least an incomplete one. "Can you rebuild it from scratch?" tests implementation recall. But look at what just happened with uroboro: I demonstrated understanding of the system not by rebuilding it from memory, but by using it, hitting its limits, and knowing exactly what was missing. The understanding showed up in use, not in an exam.
If the redistribution argument is real — if the valuable work is intent and design, not syntax and convention — then the ability to hand-code a compositing algorithm from memory isn't the thing worth testing. What matters is whether you understand why the mask expansion works the way it does, what would break if you changed it, and when you'd choose a different approach entirely.
The sphinx needs to evolve. Not "rebuild it without tools" but "explain why it works this way and what would fail if it didn't." Not testing whether you've memorized the implementation. Testing whether you understand the territory.
That's what the sphinx becomes as an uroboro hook — not a hand-coding exam, but a boundary map. You finish a session, the hook fires, it generates questions based on what you just built. Not "rewrite this function from memory" but "what assumptions does this architecture make, and which ones are load-bearing?" The conscience shows up in the workflow, not as homework in a separate repo. And it maps where your understanding ends and the AI's pattern-matching begins. Not as an anxiety exercise. As orientation.
The approval rate
I built keygrave as a satirical shitpost — a fake corporate telemetry enrollment screen that parses your actual Claude Code session data and outputs real numbers dressed in dystopian aesthetics. The three-key keyboard is the punchline: if you're pressing "approve" 94% of the time, why do you own 104 keys?
The numbers: 29,038 approvals. Zero rejections. 94.3% of my inputs are a single keystroke. Average time between Claude's output and my next input: 7.6 seconds.
The old version of this essay spent a lot of time agonizing about whether those numbers mean I've surrendered agency. I built the satire, the satire described me, existential crisis. It was a good bit. But the agonizing was more interesting than the conclusion, which tells you the conclusion was wrong.
Here's what the numbers actually describe: a permission system's granularity. The 29,009 auto-accepts are the system asking "run grep?" and me pressing enter. That's not a decision point. That's friction left over from a trust model that hasn't caught up with the workflow. Keygrave measures the friction. The decisions happened earlier — in the uroboro captures, in the architecture discussions, in the moment where I said "SQLite, not Postgres" and the AI said "okay" and built accordingly.
The 7.6 seconds is more interesting. What happens in 7.6 seconds? Not a code review — you can't review a diff in 7.6 seconds. What you can do is check whether the output matches the intent you already established. Does this look like what we discussed? Is it going in the direction I set? That's a different kind of review. It's not line-by-line verification. It's pattern-matching against a mental model. And if the mental model is well-calibrated, 7.6 seconds is enough.
If.
The question isn't whether 94.3% approval is too high. The question is whether my mental model of what the AI is doing is accurate — whether I'd notice when it drifts. Keygrave can't measure that. Neither can I, from the inside. But the sphinx can test it, if it asks the right questions.
The auteur question
Four lenses on the same nine months. The commits show a hype cycle with tools surviving the trough. The uroboro captures show intent persisting across the cycle — consistent architectural preferences regardless of which phase the commit counts are in. The sphinx shows the instinct to verify understanding, evolving from implementation recall toward boundary mapping. The approval rate shows high trust at the execution layer and a 7.6-second feedback loop that's either efficient calibration or insufficient scrutiny.
The redistribution story holds up. The intent layer — what to build, which trade-offs, which patterns — stayed human throughout. It got better-instrumented over time, not thinner. The implementation layer moved almost entirely to the AI, and the evidence suggests that's fine, because building software is largely syntax and convention. The hard part is knowing what people want when they don't know what they want.
But there's a harder question underneath.
The uroboro log shows consistent preferences across 54 projects. Zero dependencies. SQLite everywhere. CLI-first. These recur regardless of language, domain, or client. They're mine. And as the AI gets better at anticipating them — as it learns "this person always picks SQLite" — those preferences get reinforced, not examined.
Does my preference matter if it results in worse code? Sometimes Postgres is the right call. Sometimes a web UI is the right call. Sometimes a dependency saves three weeks. If the AI is following my preferences instead of optimizing for the problem, my architectural consistency isn't a strength. It's a blind spot with a philosophy attached.
Is preference in the absence of rational justification a good trait in a professional software engineer? The conventional answer is no. Engineering is supposed to be evidence-based. You pick the tool that fits the job, not the tool you always pick. Preference without justification is bias. Bias produces worse outcomes. This is the argument against auteurs in software — that personal style is an indulgence the codebase pays for.
But.
We don't actually build software that way. We build it with humans who have taste, opinions, scars from past projects, aesthetic sensibilities about what clean code looks like. A codebase with no auteur — optimized purely for local decisions, each tool chosen in isolation for each task — has no coherence. No throughline. No voice. It works, but nobody can reason about it as a whole because every part was decided independently.
The zero-dependency preference creates a body of work where any piece can be understood in isolation. The SQLite preference creates portable, self-contained systems. The CLI-first preference creates composable tools that work with the Unix philosophy and, as it turns out, work even better when the operator is a large language model. These aren't optimal local decisions. They're a coherent worldview applied across projects. The worldview produces constraints, and the constraints produce a recognizable style, and the style produces systems that fit together in ways that locally-optimized choices wouldn't.
Do we want auteurs or builders? We want both. The AI is the builder — it optimizes locally, respects the syntax, follows convention, implements cleanly. The human is the auteur — applies taste, maintains coherence, chooses constraints that create a body of work rather than a collection of isolated solutions. Neither role works alone. The builder without an auteur produces technically correct software that nobody can navigate. The auteur without a builder produces beautiful blueprints that never ship.
The redistribution put each role where it fits. The risk isn't dependency. The risk is the auteur mistaking stubbornness for taste — and the builder being too agreeable to push back.
The sphinx should be testing for that, too. Not just "do you understand what you built?" but "is your preference actually serving the problem, or is the problem being forced to serve your preference?" That's the question I don't have a tool for yet.