← Back to Distilling style

Under the hood

How the extraction pipeline actually works

February 2026

Both tools — uroboro distill and uroboro prompt-profile — follow the same three-stage architecture: extract structured data from local sources, output JSONL, and optionally feed that JSONL to an LLM for synthesis. All source code lives in the uroboro repo under internal/distill/ and internal/promptprofile/.

Extract
git log, SQLite, session JSONL
→
Filter
regex, classification, noise strip
→
Correlate
±30min window
→
JSONL
one record per line
→
Synthesize
LLM analysis prompt

The commit filter

Not all commits carry style signal. The distill command uses a 24-term extended regex to find commits where you made deliberate quality choices — refactoring, simplifying, restructuring. The regex is applied case-insensitive via git log --extended-regexp --grep.

refactor explicit restructuring
clean.?up "cleanup", "clean-up", "clean up"
simplif "simplify", "simplified", "simplification"
extract function/method extraction
rename naming improvements
reorganiz "reorganize", "reorganized"
restructur "restructure", "restructured"
consolidat merging related code
dedup "deduplicate", "dedup"
inline|flatten|split structural changes
decompos|modular "decompose", "modularize"
move.*to|pull.*out "move X to Y", "pull out Z"
reduce.*complex "reduce complexity"
improve.*read "improve readability"
untangle|decouple reducing coupling
encapsulat|abstract boundary creation
normalize consistency passes

After the regex pre-filter, two hard limits apply. Commits touching more than 20 files are dropped — these are typically bulk operations, not deliberate style choices. Diffs are capped at 50KB and truncated if exceeded. Merge commits are excluded via --no-merges.

Design choice

The regex targets the commit message, not the diff content. The assumption: if you described what you were doing with words like "refactor" or "simplify," you were making a deliberate aesthetic choice. The diff between before and after in that commit is a direct statement about what you consider better code.

The output schema

Both extraction pipelines produce JSONL — one JSON object per line. Two record types.

GitExtract

FieldTypeNotes
sourcestringalways "git"
repostringbasename of repo path
hashstringfull commit SHA
parent_hashstringfirst parent only
messagestringcommit subject line
datetimeauthor date, RFC3339
files[]stringfrom git show --numstat
languagestringplurality vote by extension
diffstringunified diff, capped at 50KB
diff_statsobject{additions, deletions, files}

UroExtract

FieldTypeNotes
sourcestringalways "uroboro"
typestringdecision | blocker | question | capture
contentstringraw capture text
projectstringfrom DB project field
tags[]stringparsed from comma-separated
timestamptimecapture timestamp
correlated_git_hashstringset by Correlate(), omitted if empty

Language detection maps file extensions to language names. The commit's language is determined by plurality vote — whichever language has the most files in the commit wins. 17 languages are recognized: Go, Python, TypeScript, JavaScript, Rust, Ruby, Java, C, C++, C#, Shell, SQL, HTML, CSS, YAML, JSON, TOML.

Capture filtering

Decisions, blockers, and questions always pass through — they're intentional captures that inherently carry signal. But generic capture-type records are filtered. They only pass if their content contains at least one of these code-related keywords:

architecture refactor implement interface schema database api struct class function module package endpoint middleware handler pattern abstraction dependency migration query index cache config deploy build test lint format convention style naming error handling

This prevents general work-log captures from diluting the style signal. A capture like "shipped the feature to staging" gets filtered out. A capture like "extracted handler to middleware for reuse" passes.

The correlation algorithm

The correlation step is what makes the combined data more valuable than either source alone. It links uro captures (intent) with git commits (implementation) when they happen within 30 minutes of each other.

// Simplified from internal/distill/correlate.go
const CorrelationWindow = 30 * time.Minute

func Correlate(git []GitExtract, uro []UroExtract) {
    for i := range uro {
        bestHash, bestDelta := "", CorrelationWindow + 1
        for _, g := range git {
            delta := abs(uro[i].Timestamp.Sub(g.Date))
            if delta <= CorrelationWindow && delta < bestDelta {
                bestDelta = delta
                bestHash = g.Hash
            }
        }
        uro[i].CorrelatedGitHash = bestHash
    }
}

Key properties:

Bidirectional
The capture can come before or after the commit. "I'm going to restructure this" (capture) followed by the restructuring commit, or the commit followed by "decided to restructure because..." (decision capture). Both correlate.
Nearest-wins
If multiple commits fall within the 30-minute window, the closest one wins. This prevents a capture at 14:00 from correlating with a commit at 13:31 when there's a closer commit at 13:55.
Project-blind
Correlation ignores whether the uro capture's project tag matches the git repo. It's purely temporal. This is acknowledged as a simplification — at milestone 1 scale (hundreds of records, not millions), false correlations are rare enough to not warrant the complexity of project matching.
O(n×m)
A full cross-product scan. Every uro capture is compared against every git commit. The code comment acknowledges this and calls it "fine at milestone 1 scale." For the typical extraction (50–100 git commits, 100–500 uro captures), this completes in milliseconds.
The signal

A correlated pair gives you both the "why" (from the uro decision) and the "what" (from the git diff). The analysis prompt calls these "the richest signal" — when you have a decision record saying "guard clauses over nested ifs for readability" paired with a diff that shows the actual before/after transformation, you have a concrete, evidenced style rule.

Prompt classification

The prompt-profile tool classifies each user message along four dimensions: imperative, question, file reference, and code block. The detection is heuristic — first-word matching, not semantic analysis.

Imperative detection

Two-stage check. First, the lowercase first word must match one of 57 imperative verbs:

fix add implement create update remove delete refactor change make build write run test check move rename extract clean simplify merge deploy configure setup install upgrade migrate convert replace review debug optimize search find list show help explain design plan analyze summarize generate format validate read open start stop set get use try apply commit push pull revert reset look

Second check: the prompt starts with one of 10 directive phrases:

let's can you could you please go ahead we need to we should i want to i need i'd like
Edge case

"Can you fix this?" registers as both imperative (via "can you" directive phrase) and question (via "can" question starter + "?"). This overlap is intentional — some messages genuinely are both a request and a question. The prompt anatomy data shows this dual-classification accounts for much of the "other" category.

Question detection

Also two-stage. Contains ? anywhere in the text, or the first word matches one of 18 question starters:

how what why when where which can could should would is are do does did will has have

File path detection

Word-by-word scan. A word is a file path if it starts with / (and contains another /), starts with ./ or ../, or contains a dot where the part after the last dot matches one of 30+ recognized extensions: .go, .py, .ts, .tsx, .js, .jsx, .rs, .html, .css, .json, .yaml, .md, .sql, .sh, and more. URLs starting with http are excluded.

Noise stripping

Before classification, the tool strips system-injected content from the raw session JSONL. XML-wrapped blocks like <system-reminder>, <task-notification>, <command-name>, and <bash-notification> are removed. Messages starting with known system prefixes are dropped entirely. Sidechain sessions (subagent conversations) are skipped based on the IsSidechain flag in the session index.

The analysis prompt

The included scripts/style-analysis-prompt.md defines a 6-step workflow for turning JSONL into a style profile. It's designed to be fed to Claude along with the extracted data.

Step 1 — Inventory
Count records by source, language, repo. Establish what data is available before analyzing.
Step 2 — Analyze git extracts
Look at diff + commit message together for: naming patterns, structure preferences (guard clauses vs nesting, function length, file organization), error handling, abstraction thresholds, simplification patterns, testing patterns. Group by language.
Step 3 — Analyze uro extracts
Parse the "X over Y — reason" format to extract preferred approach, rejected approach, and reasoning. Group decisions by category: architecture, tooling, code style, dependencies, testing, workflow.
Step 4 — Cross-reference correlated pairs
For each uro extract with a correlated_git_hash, pair it with its git extract. Intent meets implementation. The prompt calls these "the richest signal."
Step 5 — Cluster and rank
Rules ranked by: frequency, breadth (repos/languages), intent (backed by uro decision), consistency. 8+ occurrences across 3 repos with uro backing = hard rule. 1 commit, no uro context = candidate.
Step 6 — Produce three outputs
Style guide, system prompt fragment, and machine-readable rules JSON.

The three output artifacts:

STYLE_GUIDE.md

Full reference with before/after code from your actual commits. Uses diff lines (- = before, + = after) as examples.

~2000 words

STYLE_PROMPT.md

System prompt fragment with hard rules, preferences, and anti-patterns. Ready to paste into CLAUDE.md.

<1500 tokens

style_rules.json

Machine-readable rules with fields: id, rule, category, severity, languages, frequency, repos, example_hash, rationale.

structured data
Key instruction

The prompt explicitly says: "Use actual code snippets from the diffs for before/after examples. Don't invent synthetic examples." The style guide's credibility depends on every example being traceable to a real commit hash.

Multi-repo extraction

The batch script scripts/distill-multi.sh orchestrates extraction across multiple repositories. It runs uroboro distill --source git for each repo in sequence, appending to one JSONL file. Uro captures are extracted once globally (they're not repo-scoped). The output includes a manifest tracking provenance.

$ ./scripts/distill-multi.sh --days 90 --correlate \
    ~/projects/uroboro ~/projects/qryzone ~/projects/sjiek

extracting git from uroboro... 47 commits
extracting git from qryzone... 12 commits
extracting git from sjiek... 8 commits
extracting uro captures... 83 records
correlating (30min window)... 12 pairs
wrote extracts-2026-02-17.jsonl (150 records)
wrote extract-manifest.json

The manifest is a provenance record:

{
  "date": "2026-02-17",
  "repos": ["uroboro", "qryzone", "sjiek"],
  "git_count": 67,
  "uro_count": 83,
  "correlated_count": 12,
  "total_records": 150,
  "days": 90,
  "output": "extracts-2026-02-17.jsonl"
}

When invoked via MCP (the uro_distill tool), output goes to ~/.local/share/uroboro/style-data/ with timestamped filenames. Same JSONL format, same schema — the MCP layer is a thin wrapper over the same Go code.

Known limitations

The correlation is project-blind. A capture tagged "qryzone" can correlate with a commit in a different repo if the timestamps are close enough. For the current scale this rarely produces false matches, but it's a known simplification.

The batch script's --correlate mode only correlates uro captures against the first repo in the list, not all repos. This is acknowledged in a source comment as a "simplification for now."

The imperative classifier misses common patterns. "Great, now add routing" starts with "great," which isn't a verb, so the imperative is missed. Compound messages with multiple intents get classified by their first word only. The 42% "other" category in the prompt anatomy data reflects these classification gaps.

No test fixtures exist for the distill package. The extraction logic has been validated empirically (by running it against real repos and checking the output), but there are no automated regression tests for the commit filtering or correlation logic.