Both tools — uroboro distill and uroboro prompt-profile — follow the same three-stage architecture: extract structured data from local sources, output JSONL, and optionally feed that JSONL to an LLM for synthesis. All source code lives in the uroboro repo under internal/distill/ and internal/promptprofile/.
The commit filter
Not all commits carry style signal. The distill command uses a 24-term extended regex to find commits where you made deliberate quality choices — refactoring, simplifying, restructuring. The regex is applied case-insensitive via git log --extended-regexp --grep.
After the regex pre-filter, two hard limits apply. Commits touching more than 20 files are dropped — these are typically bulk operations, not deliberate style choices. Diffs are capped at 50KB and truncated if exceeded. Merge commits are excluded via --no-merges.
The regex targets the commit message, not the diff content. The assumption: if you described what you were doing with words like "refactor" or "simplify," you were making a deliberate aesthetic choice. The diff between before and after in that commit is a direct statement about what you consider better code.
The output schema
Both extraction pipelines produce JSONL — one JSON object per line. Two record types.
GitExtract
| Field | Type | Notes |
|---|---|---|
| source | string | always "git" |
| repo | string | basename of repo path |
| hash | string | full commit SHA |
| parent_hash | string | first parent only |
| message | string | commit subject line |
| date | time | author date, RFC3339 |
| files | []string | from git show --numstat |
| language | string | plurality vote by extension |
| diff | string | unified diff, capped at 50KB |
| diff_stats | object | {additions, deletions, files} |
UroExtract
| Field | Type | Notes |
|---|---|---|
| source | string | always "uroboro" |
| type | string | decision | blocker | question | capture |
| content | string | raw capture text |
| project | string | from DB project field |
| tags | []string | parsed from comma-separated |
| timestamp | time | capture timestamp |
| correlated_git_hash | string | set by Correlate(), omitted if empty |
Language detection maps file extensions to language names. The commit's language is determined by plurality vote — whichever language has the most files in the commit wins. 17 languages are recognized: Go, Python, TypeScript, JavaScript, Rust, Ruby, Java, C, C++, C#, Shell, SQL, HTML, CSS, YAML, JSON, TOML.
Capture filtering
Decisions, blockers, and questions always pass through — they're intentional captures that inherently carry signal. But generic capture-type records are filtered. They only pass if their content contains at least one of these code-related keywords:
This prevents general work-log captures from diluting the style signal. A capture like "shipped the feature to staging" gets filtered out. A capture like "extracted handler to middleware for reuse" passes.
The correlation algorithm
The correlation step is what makes the combined data more valuable than either source alone. It links uro captures (intent) with git commits (implementation) when they happen within 30 minutes of each other.
// Simplified from internal/distill/correlate.go
const CorrelationWindow = 30 * time.Minute
func Correlate(git []GitExtract, uro []UroExtract) {
for i := range uro {
bestHash, bestDelta := "", CorrelationWindow + 1
for _, g := range git {
delta := abs(uro[i].Timestamp.Sub(g.Date))
if delta <= CorrelationWindow && delta < bestDelta {
bestDelta = delta
bestHash = g.Hash
}
}
uro[i].CorrelatedGitHash = bestHash
}
}
Key properties:
A correlated pair gives you both the "why" (from the uro decision) and the "what" (from the git diff). The analysis prompt calls these "the richest signal" — when you have a decision record saying "guard clauses over nested ifs for readability" paired with a diff that shows the actual before/after transformation, you have a concrete, evidenced style rule.
Prompt classification
The prompt-profile tool classifies each user message along four dimensions: imperative, question, file reference, and code block. The detection is heuristic — first-word matching, not semantic analysis.
Imperative detection
Two-stage check. First, the lowercase first word must match one of 57 imperative verbs:
Second check: the prompt starts with one of 10 directive phrases:
"Can you fix this?" registers as both imperative (via "can you" directive phrase) and question (via "can" question starter + "?"). This overlap is intentional — some messages genuinely are both a request and a question. The prompt anatomy data shows this dual-classification accounts for much of the "other" category.
Question detection
Also two-stage. Contains ? anywhere in the text, or the first word matches one of 18 question starters:
File path detection
Word-by-word scan. A word is a file path if it starts with / (and contains another /), starts with ./ or ../, or contains a dot where the part after the last dot matches one of 30+ recognized extensions: .go, .py, .ts, .tsx, .js, .jsx, .rs, .html, .css, .json, .yaml, .md, .sql, .sh, and more. URLs starting with http are excluded.
Noise stripping
Before classification, the tool strips system-injected content from the raw session JSONL. XML-wrapped blocks like <system-reminder>, <task-notification>, <command-name>, and <bash-notification> are removed. Messages starting with known system prefixes are dropped entirely. Sidechain sessions (subagent conversations) are skipped based on the IsSidechain flag in the session index.
The analysis prompt
The included scripts/style-analysis-prompt.md defines a 6-step workflow for turning JSONL into a style profile. It's designed to be fed to Claude along with the extracted data.
correlated_git_hash, pair it with its git extract. Intent meets implementation. The prompt calls these "the richest signal."The three output artifacts:
STYLE_GUIDE.md
Full reference with before/after code from your actual commits. Uses diff lines (- = before, + = after) as examples.
STYLE_PROMPT.md
System prompt fragment with hard rules, preferences, and anti-patterns. Ready to paste into CLAUDE.md.
style_rules.json
Machine-readable rules with fields: id, rule, category, severity, languages, frequency, repos, example_hash, rationale.
The prompt explicitly says: "Use actual code snippets from the diffs for before/after examples. Don't invent synthetic examples." The style guide's credibility depends on every example being traceable to a real commit hash.
Multi-repo extraction
The batch script scripts/distill-multi.sh orchestrates extraction across multiple repositories. It runs uroboro distill --source git for each repo in sequence, appending to one JSONL file. Uro captures are extracted once globally (they're not repo-scoped). The output includes a manifest tracking provenance.
$ ./scripts/distill-multi.sh --days 90 --correlate \
~/projects/uroboro ~/projects/qryzone ~/projects/sjiek
extracting git from uroboro... 47 commits
extracting git from qryzone... 12 commits
extracting git from sjiek... 8 commits
extracting uro captures... 83 records
correlating (30min window)... 12 pairs
wrote extracts-2026-02-17.jsonl (150 records)
wrote extract-manifest.json
The manifest is a provenance record:
{
"date": "2026-02-17",
"repos": ["uroboro", "qryzone", "sjiek"],
"git_count": 67,
"uro_count": 83,
"correlated_count": 12,
"total_records": 150,
"days": 90,
"output": "extracts-2026-02-17.jsonl"
}
When invoked via MCP (the uro_distill tool), output goes to ~/.local/share/uroboro/style-data/ with timestamped filenames. Same JSONL format, same schema — the MCP layer is a thin wrapper over the same Go code.
Known limitations
The correlation is project-blind. A capture tagged "qryzone" can correlate with a commit in a different repo if the timestamps are close enough. For the current scale this rarely produces false matches, but it's a known simplification.
The batch script's --correlate mode only correlates uro captures against the first repo in the list, not all repos. This is acknowledged in a source comment as a "simplification for now."
The imperative classifier misses common patterns. "Great, now add routing" starts with "great," which isn't a verb, so the imperative is missed. Compound messages with multiple intents get classified by their first word only. The 42% "other" category in the prompt anatomy data reflects these classification gaps.
No test fixtures exist for the distill package. The extraction logic has been validated empirically (by running it against real repos and checking the output), but there are no automated regression tests for the commit filtering or correlation logic.