← Back to Explore seedling

Distilling style

Mining git history and AI conversations to extract how you actually work

February 2026

You write code for months. Years. Decisions accumulate. You start preferring guard clauses over nested ifs. You extract functions at a certain threshold. You name things a certain way, handle errors a certain way, structure files a certain way. These patterns stay implicit — lodged in muscle memory and scattered across thousands of commits. You couldn't write them down if someone asked.

Same thing happens with how you talk to AI. You develop prompting habits. Certain sentence structures. A ratio of questions to commands. Times of day when you're most active. None of it deliberate. All of it real.

What if you could extract these patterns from evidence? Not introspect and guess — actually mine the data?

Two tools

I built two extraction commands into uroboro. The first, distill, mines git history and uroboro captures for code style signals. The second, prompt-profile, mines Claude Code conversation logs for prompting patterns. Both produce structured JSONL that can be fed to an LLM for synthesis, or analyzed directly.

They answer different questions. Distill asks: how do you write code? Prompt-profile asks: how do you talk to AI? Together they build a picture of how you actually work — from evidence, not self-report.

Discovering your code style

The distill command looks for commits where you cared about code quality. Not every commit — specifically the ones where you refactored, simplified, extracted, renamed, restructured. The signal commits. It matches against commit messages using a regex that catches words like refactor, simplify, extract, consolidate, decompose, decouple. If you bothered to describe what you were doing with those words, you were making a deliberate style choice.

For each matching commit, it extracts the full unified diff (capped at 50KB), the files touched, a language detection from file extensions, and diff stats. It also pulls in any uroboro captures you made — decisions, blockers, questions — and optionally correlates them: if a capture was recorded within 30 minutes of a git commit, they're linked. Intent meets implementation.

Run it across multiple repos with the batch script and you get a consolidated JSONL file covering your style across projects. Feed that to Claude with the included analysis prompt and it produces three outputs: a style guide with before/after examples from your own commits, a system prompt fragment with hard rules and preferences, and a machine-readable style_rules.json with severity ratings.

$ uroboro distill --source all --correlate --repo ~/projects/uroboro
extracted 47 git commits (style-signal)
extracted 83 uro captures (code-related)
correlated 12 pairs (30min window)
total: 130 records → stdout

47 commits out of however many hundreds. That filtering matters. If you fed your entire git history to an LLM and asked for style patterns, you'd get noise. The commits where you explicitly cleaned up, restructured, simplified — those are the ones where your aesthetic preferences are visible. The diff between "before" and "after" in a refactoring commit is a direct statement about what you consider better.

Prompting introspection

Prompt-profile does something different. It reads every Claude Code session on your machine — the JSONL files that record each conversation — and extracts every user prompt. Then it classifies them. Is this an imperative ("fix the auth bug")? A question ("why is this returning nil?")? Does it reference a file path? Does it contain a code block? It strips out injected system content, filters noise, and computes stats across your entire usage history.

$ uroboro prompt-profile --stats

Prompt Profile Analysis
=======================

Overview
  Sessions:     125
  Prompts:      1,106
  Words:        38,710
  Period:       2025-06-03 → 2026-02-15

Length Distribution
  Short  (<20w):     482  43.6%  ████████████████████
  Medium (20-100w):  438  39.6%  ██████████████████
  Long   (100-500w): 162  14.6%  ██████
  Very long (500w+):  24   2.2%  █

Style
  Imperative:  561  50.7%
  Questions:   465  42.0%
  File refs:   287  25.9%
  Code blocks: 104   9.4%

Active Times
  Peak hour: 14:00 (Tuesday)
  Peak day:  Tuesday

There it is. 1,106 prompts since June 2025. Nearly half are questions. My prompts are overwhelmingly short — 43% under 20 words. I reference file paths in a quarter of messages. I peak on Tuesdays at 14:00. None of this was intentional. All of it is pattern.

The 42% question rate was the surprise. I thought I mostly gave commands. Turns out I ask nearly as often as I direct. "Why is this failing?" and "what does this function do?" almost as frequent as "fix this" and "add that." The data says I use Claude as much for understanding as for execution. That changes how you'd optimize a system prompt for my usage — it needs to be good at explanation, not just task completion.

Building a developer profile

The real use case is combining both outputs. Your code style distillation tells a sub-agent how you write code — naming conventions, error handling patterns, when you extract functions, how you structure files. Your prompt profile tells it how you communicate — short directives, frequent questions, file path references. Together they form a developer profile that can be injected as system prompt context.

This matters for sub-agents. When Claude Code spawns a task agent to handle something in the background, that agent starts from zero. It doesn't know you prefer guard clauses, that you use fmt.Errorf with %w wrapping, that your functions tend to be under 30 lines, that you name test files *_test.go with table-driven tests. A distilled style prompt gives the sub-agent your preferences without you restating them every session.

Same for the prompting profile. If a system knows you tend to ask short questions and expect concise answers, it can calibrate response length. If it knows you reference file paths frequently, it can prioritize code context over general explanation. The profile becomes a calibration layer.

The data pipeline

Three stages. Extraction, analysis, synthesis.

Extraction is the mechanical part. Git log with regex filtering produces style-signal commits. The uroboro database produces decisions, blockers, questions — the intentional captures. Prompt-profile reads Claude Code session files and classifies every user message. All outputs are JSONL: one JSON object per line, streamable, composable.

$ ./scripts/distill-multi.sh --days 90 --correlate \
    ~/projects/uroboro ~/projects/qryzone ~/projects/sjiek

extracting git from uroboro... 47 commits
extracting git from qryzone... 12 commits
extracting git from sjiek... 8 commits
extracting uro captures... 83 records
correlating (30min window)... 12 pairs
wrote extracts-2026-02-17.jsonl (150 records)
wrote manifest-2026-02-17.json

The batch script runs distill across multiple repos, extracts uro captures once globally, and writes a consolidated file with a manifest tracking provenance: which repos, how many records, how many correlations.

Analysis is where the LLM comes in. The included style-analysis-prompt.md instructs Claude to inventory the records, analyze naming patterns, structure preferences, error handling, abstraction thresholds, and testing patterns from the git extracts. It cross-references with uro captures to find places where intent (the decision record) aligns with implementation (the commit diff). Correlated pairs are the strongest signal — you said why you were doing something, and the diff shows what you did.

Synthesis produces three artifacts. A style guide (~2000 words) with before/after code snippets pulled from your actual commits. A system prompt fragment (<1500 tokens) with hard rules, preferences, and anti-patterns. And a machine-readable style_rules.json with severity, frequency, and example commit hashes for each rule. The prompt fragment is the one you'd actually use day-to-day.

What it revealed

125 sessions. 1,106 prompts. 38,710 words. Eight months of working with Claude Code, reduced to a statistical profile.

The length distribution is lopsided toward brevity. 43.6% of my prompts are under 20 words. Another 39.6% are 20-100 words. Only 2.2% exceed 500 words. I'm terse. When I write long prompts, they tend to contain code blocks or detailed plans — the 9.4% code block rate correlates with the longer messages. My default mode is short directive or short question.

50.7% imperative, 42% questions. Nearly balanced. The imperative detection checks for verbs like fix, add, implement, refactor, run, test at the start of a message, plus directive phrases like "can you" and "let's." The question detection looks for question marks and question-starting words. Some messages are both — "can you fix this?" registers as both imperative and question. The near-balance means I alternate between directing and understanding. Not one mode. A conversation.

25.9% of prompts reference file paths. That's high. It means I'm constantly pointing at specific locations in the codebase rather than describing things abstractly. "Look at internal/distill/git.go" rather than "look at the git extraction code." Direct reference over description. This tracks with the short prompt length — a file path does a lot of contextual work in few characters.

Peak Tuesdays at 14:00. After lunch. I have no explanation for this. Maybe Monday is for orientation and Tuesday is when actual building happens. Maybe it's arbitrary. The data says what it says.

The numbers, visualized

Sessions
125
8 months
Prompts
1,106
~9/session
Words
38,710
~35/prompt
Peak
Tue 14h
after lunch
Prompt length
<20 words
482 44%
20–100
438 40%
100–500
162 15%
500+
24 2%
Style split
51% imperative 42% questions
Context signals
File paths 26%
Code blocks 9%
Weekly activity
Mon
Tue
Wed
Thu
Fri
Sat
Sun

What this is for

The obvious use: better system prompts. If you're running Claude Code with a CLAUDE.md, your distilled style guide gives you rules grounded in your actual behavior rather than aspirational preferences. You don't have to guess what you prefer. The evidence is in the commits.

The less obvious use: self-knowledge. I didn't know I asked questions 42% of the time. I didn't know my prompts were that short. I didn't know I peaked on Tuesdays. These patterns were invisible until I extracted them. Now I can decide if they're what I want, or if they're habits worth changing.

The speculative use: cross-project consistency. Run distill across all your repos and the style guide reflects your actual cross-project patterns, not the conventions of any single codebase. If you consistently use early returns across Go, Python, and TypeScript, that shows up. If you handle errors differently per language, that shows up too. The tool doesn't prescribe — it describes.

All of this runs locally. No data leaves your machine unless you choose to send the JSONL to an API. The extraction reads files on disk — git repos, SQLite databases, Claude Code session logs. The analysis can run through a local model if you want. The same local-first constraint that shaped uroboro's architecture applies here.

The deep dive

Prompt anatomy What 1,106 prompts reveal Code decisions 534 captures, 8 patterns Under the hood The extraction pipeline