Distilling style
Mining git history and AI conversations to extract how you actually work
February 2026
You write code for months. Years. Decisions accumulate. You start preferring guard clauses over nested ifs. You extract functions at a certain threshold. You name things a certain way, handle errors a certain way, structure files a certain way. These patterns stay implicit — lodged in muscle memory and scattered across thousands of commits. You couldn't write them down if someone asked.
Same thing happens with how you talk to AI. You develop prompting habits. Certain sentence structures. A ratio of questions to commands. Times of day when you're most active. None of it deliberate. All of it real.
What if you could extract these patterns from evidence? Not introspect and guess — actually mine the data?
Two tools
I built two extraction commands into uroboro. The first, distill, mines git history and uroboro captures for code style signals. The second, prompt-profile, mines Claude Code conversation logs for prompting patterns. Both produce structured JSONL that can be fed to an LLM for synthesis, or analyzed directly.
They answer different questions. Distill asks: how do you write code? Prompt-profile asks: how do you talk to AI? Together they build a picture of how you actually work — from evidence, not self-report.
Discovering your code style
The distill command looks for commits where you cared about code quality. Not every commit — specifically the ones where you refactored, simplified, extracted, renamed, restructured. The signal commits. It matches against commit messages using a regex that catches words like refactor, simplify, extract, consolidate, decompose, decouple. If you bothered to describe what you were doing with those words, you were making a deliberate style choice.
For each matching commit, it extracts the full unified diff (capped at 50KB), the files touched, a language detection from file extensions, and diff stats. It also pulls in any uroboro captures you made — decisions, blockers, questions — and optionally correlates them: if a capture was recorded within 30 minutes of a git commit, they're linked. Intent meets implementation.
Run it across multiple repos with the batch script and you get a consolidated JSONL file covering your style across projects. Feed that to Claude with the included analysis prompt and it produces three outputs: a style guide with before/after examples from your own commits, a system prompt fragment with hard rules and preferences, and a machine-readable style_rules.json with severity ratings.
$ uroboro distill --source all --correlate --repo ~/projects/uroboro
extracted 47 git commits (style-signal)
extracted 83 uro captures (code-related)
correlated 12 pairs (30min window)
total: 130 records → stdout
47 commits out of however many hundreds. That filtering matters. If you fed your entire git history to an LLM and asked for style patterns, you'd get noise. The commits where you explicitly cleaned up, restructured, simplified — those are the ones where your aesthetic preferences are visible. The diff between "before" and "after" in a refactoring commit is a direct statement about what you consider better.
Prompting introspection
Prompt-profile does something different. It reads every Claude Code session on your machine — the JSONL files that record each conversation — and extracts every user prompt. Then it classifies them. Is this an imperative ("fix the auth bug")? A question ("why is this returning nil?")? Does it reference a file path? Does it contain a code block? It strips out injected system content, filters noise, and computes stats across your entire usage history.
$ uroboro prompt-profile --stats
Prompt Profile Analysis
=======================
Overview
Sessions: 125
Prompts: 1,106
Words: 38,710
Period: 2025-06-03 → 2026-02-15
Length Distribution
Short (<20w): 482 43.6% ████████████████████
Medium (20-100w): 438 39.6% ██████████████████
Long (100-500w): 162 14.6% ██████
Very long (500w+): 24 2.2% █
Style
Imperative: 561 50.7%
Questions: 465 42.0%
File refs: 287 25.9%
Code blocks: 104 9.4%
Active Times
Peak hour: 14:00 (Tuesday)
Peak day: Tuesday
There it is. 1,106 prompts since June 2025. Nearly half are questions. My prompts are overwhelmingly short — 43% under 20 words. I reference file paths in a quarter of messages. I peak on Tuesdays at 14:00. None of this was intentional. All of it is pattern.
The 42% question rate was the surprise. I thought I mostly gave commands. Turns out I ask nearly as often as I direct. "Why is this failing?" and "what does this function do?" almost as frequent as "fix this" and "add that." The data says I use Claude as much for understanding as for execution. That changes how you'd optimize a system prompt for my usage — it needs to be good at explanation, not just task completion.
Building a developer profile
The real use case is combining both outputs. Your code style distillation tells a sub-agent how you write code — naming conventions, error handling patterns, when you extract functions, how you structure files. Your prompt profile tells it how you communicate — short directives, frequent questions, file path references. Together they form a developer profile that can be injected as system prompt context.
This matters for sub-agents. When Claude Code spawns a task agent to handle something in the background, that agent starts from zero. It doesn't know you prefer guard clauses, that you use fmt.Errorf with %w wrapping, that your functions tend to be under 30 lines, that you name test files *_test.go with table-driven tests. A distilled style prompt gives the sub-agent your preferences without you restating them every session.
Same for the prompting profile. If a system knows you tend to ask short questions and expect concise answers, it can calibrate response length. If it knows you reference file paths frequently, it can prioritize code context over general explanation. The profile becomes a calibration layer.
The data pipeline
Three stages. Extraction, analysis, synthesis.
Extraction is the mechanical part. Git log with regex filtering produces style-signal commits. The uroboro database produces decisions, blockers, questions — the intentional captures. Prompt-profile reads Claude Code session files and classifies every user message. All outputs are JSONL: one JSON object per line, streamable, composable.
$ ./scripts/distill-multi.sh --days 90 --correlate \
~/projects/uroboro ~/projects/qryzone ~/projects/sjiek
extracting git from uroboro... 47 commits
extracting git from qryzone... 12 commits
extracting git from sjiek... 8 commits
extracting uro captures... 83 records
correlating (30min window)... 12 pairs
wrote extracts-2026-02-17.jsonl (150 records)
wrote manifest-2026-02-17.json
The batch script runs distill across multiple repos, extracts uro captures once globally, and writes a consolidated file with a manifest tracking provenance: which repos, how many records, how many correlations.
Analysis is where the LLM comes in. The included style-analysis-prompt.md instructs Claude to inventory the records, analyze naming patterns, structure preferences, error handling, abstraction thresholds, and testing patterns from the git extracts. It cross-references with uro captures to find places where intent (the decision record) aligns with implementation (the commit diff). Correlated pairs are the strongest signal — you said why you were doing something, and the diff shows what you did.
Synthesis produces three artifacts. A style guide (~2000 words) with before/after code snippets pulled from your actual commits. A system prompt fragment (<1500 tokens) with hard rules, preferences, and anti-patterns. And a machine-readable style_rules.json with severity, frequency, and example commit hashes for each rule. The prompt fragment is the one you'd actually use day-to-day.
What it revealed
125 sessions. 1,106 prompts. 38,710 words. Eight months of working with Claude Code, reduced to a statistical profile.
The length distribution is lopsided toward brevity. 43.6% of my prompts are under 20 words. Another 39.6% are 20-100 words. Only 2.2% exceed 500 words. I'm terse. When I write long prompts, they tend to contain code blocks or detailed plans — the 9.4% code block rate correlates with the longer messages. My default mode is short directive or short question.
50.7% imperative, 42% questions. Nearly balanced. The imperative detection checks for verbs like fix, add, implement, refactor, run, test at the start of a message, plus directive phrases like "can you" and "let's." The question detection looks for question marks and question-starting words. Some messages are both — "can you fix this?" registers as both imperative and question. The near-balance means I alternate between directing and understanding. Not one mode. A conversation.
25.9% of prompts reference file paths. That's high. It means I'm constantly pointing at specific locations in the codebase rather than describing things abstractly. "Look at internal/distill/git.go" rather than "look at the git extraction code." Direct reference over description. This tracks with the short prompt length — a file path does a lot of contextual work in few characters.
Peak Tuesdays at 14:00. After lunch. I have no explanation for this. Maybe Monday is for orientation and Tuesday is when actual building happens. Maybe it's arbitrary. The data says what it says.
The numbers, visualized
What this is for
The obvious use: better system prompts. If you're running Claude Code with a CLAUDE.md, your distilled style guide gives you rules grounded in your actual behavior rather than aspirational preferences. You don't have to guess what you prefer. The evidence is in the commits.
The less obvious use: self-knowledge. I didn't know I asked questions 42% of the time. I didn't know my prompts were that short. I didn't know I peaked on Tuesdays. These patterns were invisible until I extracted them. Now I can decide if they're what I want, or if they're habits worth changing.
The speculative use: cross-project consistency. Run distill across all your repos and the style guide reflects your actual cross-project patterns, not the conventions of any single codebase. If you consistently use early returns across Go, Python, and TypeScript, that shows up. If you handle errors differently per language, that shows up too. The tool doesn't prescribe — it describes.
All of this runs locally. No data leaves your machine unless you choose to send the JSONL to an API. The extraction reads files on disk — git repos, SQLite databases, Claude Code session logs. The analysis can run through a local model if you want. The same local-first constraint that shaped uroboro's architecture applies here.
The deep dive
Prompt anatomy What 1,106 prompts reveal Code decisions 534 captures, 8 patterns Under the hood The extraction pipeline