Convergent Evolution in Personal AI Tooling
How homebrew experiments from June 2025 anticipated the architecture of OpenClaw — and what that says about where personal AI is heading
February 2026
In late January 2026, OpenClaw shipped and went viral. A personal AI assistant you run on your own devices, routing through WhatsApp, Telegram, Slack — a gateway architecture with persistent memory, proactive scheduling, and a permission system for when the agent wants to run commands on your machine. Shortly after, Nader Dabit published a gist called "You Could've Invented OpenClaw", walking through how you'd build its architecture from first principles.
The premise is that the design is inevitable. If you think seriously about making LLMs useful as persistent collaborators, you'll arrive at the same set of solutions because you're hitting the same set of walls.
I found this uncomfortably familiar. Not because I'd built OpenClaw — I hadn't — but because I'd been solving the same problems with markdown files and Go CLIs eight months earlier, and the solutions looked disturbingly similar.
Timeline
It started with clipboard workflows. May 17, 2025: first commit on sjiek, a Go CLI that generates git diffs and copies them to your clipboard. Default output directory: ~/llm_context_diffs. The name is Dutch for "gum" — chew on this. It was built for one purpose: getting code context into ChatGPT and Claude's web interfaces, because that's how I was working with LLMs. Copy diff, switch to browser, paste, ask question, read answer, switch back to terminal. Repeat.
Two weeks later, May 30: first commit on uroboro, a tool for capturing development decisions, blockers, and context into a local SQLite database. By June 7, I'd written 750 lines of AI collaboration infrastructure in a single commit — session startup procedures, context briefings, timekeeping, north star principles. By June 8, there was a backup script preserving 76 uroboro captures from just nine days of use. By early June, wherewasi existed as a companion tool for context recovery across AI sessions.
The progression tells you something: clipboard bridge → structured capture → full session infrastructure. Each tool was born from hitting the next wall. Sjiek solved "I need to get context to the AI." Uroboro solved "I need to remember what happened across sessions." The core/ai/ procedures solved "I need discipline around how sessions start, run, and end." Each solution exposed the next problem.
November 24, 2025: OpenClaw's first commit. Late January 2026: it goes viral. February 2026: the "You Could've Invented OpenClaw" gist lands.
Eight months between these two independent efforts at solving the same structural problems. The solutions converged because the problems are the same for anyone using LLMs as daily collaborators.
The parallel map
Each of these parallels is backed by actual files from both projects.
Identity: CONTEXT_BRIEFING.md vs SOUL.md
The first problem you hit with LLMs is that every session starts from zero. The model doesn't know who you are, what you're building, or what matters to you. Both projects solved this with a markdown file that defines the AI's orientation.
My CONTEXT_BRIEFING.md (June 11, 2025) served as a master orientation document. It defined QRY as a methodology, listed the toolset, described the philosophy, and included guidelines for AI assistants — things like respecting local-first privacy and considering human psychology. It was the document you'd hand to a new AI session so it understood the territory.
OpenClaw's SOUL.md does the same thing with different priorities. Where mine was structured and methodology-heavy, theirs is personality-driven: "Be genuinely helpful, not performatively helpful. Have opinions. You're allowed to disagree." Both files are living documents the agent is expected to evolve. Both exist because a stateless model needs a persistent identity, and a markdown file is the simplest way to provide one.
I also had QRY_NORTHSTARS.md — a philosophy document defining ten core principles, from psychology-informed design to marketing honesty. OpenClaw doesn't have a direct equivalent. Their philosophy is baked into the SOUL.md template and expressed pragmatically through features. This divergence matters, and I'll come back to it.
Temporal awareness: TIMEKEEPING.md vs current-time.ts
LLMs don't know what day it is. This sounds trivial until you're trying to use one as a collaborator and it can't tell whether your deadline is tomorrow or three weeks ago. Both projects solved this, but the how is revealing.
My TIMEKEEPING.md (June 11, 2025) was a manually-updated file. At the start of each session, I'd write the current date, time, session type, and focus area. The AI would read it and know where we were temporally. Simple, effective, and entirely manual.
OpenClaw's current-time.ts does the same thing programmatically. A function appends timezone-aware time to every prompt: Current time: 2026-02-16 14:30:00 (America/New_York). It checks whether the time is already present to avoid duplication. Automated, reliable, zero friction.
Same problem, different levels of automation. Mine was a markdown file because I was one person with a text editor. Theirs is code because they're building a product. The architectural instinct is identical.
Memory: uroboro + wherewasi vs MEMORY.md + memory_search
Context loss across sessions is the big one. You make a decision on Tuesday, start a new chat on Thursday, and the model has no idea what you decided. Both projects built structured capture systems, and both ended up with remarkably similar data models.
Uroboro stores captures in SQLite: timestamp, content, project (auto-detected from git), tags (auto-enhanced from content analysis), and git branch. It has specific commands for decisions (uro d "JWT over sessions — stateless"), blockers (uro b "waiting on API docs"), and questions (uro q "revocation strategy?"). Wherewasi complements this by tracking git commits, file changes, and generating dense context summaries you can pull into a new AI session in seconds.
OpenClaw stores memories in MEMORY.md and topical files under memory/*.md, then indexes them with embedding-based semantic search. You can memory_search with a natural language query and get ranked results with citations. Sessions themselves are stored as JSONL files — append-only, durable, inspectable.
The structural parallel is clear: both systems capture context for retrieval across sessions. The design divergence is where it gets revealing. Uroboro captures deliberately — you choose to record a decision or flag a blocker, and the system enhances your input with auto-detected metadata. OpenClaw captures more broadly and relies on semantic search to find what's relevant later. Precision versus recall, in information retrieval terms.
Session lifecycle: AI_SESSION_PROCEDURE.md vs session management
My AI_SESSION_PROCEDURE.md (June 21, 2025) defined a startup/shutdown protocol: check the environment, generate a morning digest, establish session context, verify uroboro status. At shutdown: summarize accomplishments, capture key events, update the context briefing, run quality assurance. It was a checklist I followed every session to maintain discipline.
OpenClaw automates this lifecycle. Sessions are scoped per-sender with configurable reset modes (daily, idle timeout, manual). JSONL transcripts persist across turns. Heartbeats provide scheduled check-ins. The session snapshot system ensures skills and configuration stay consistent within a session but pick up changes on the next one.
Again: sessions need lifecycle management. The difference is human discipline versus automated infrastructure.
I benchmarked this overhead on June 8. A fresh morning context restoration took 13 tool calls, reading ~18,700 words of structured documentation to produce a working digest — about four minutes to get a new AI session oriented on the full project state. That's the exact startup cost both architectures are trying to eliminate. My target for wherewasi was to cut it to under five tool calls and sixty seconds. OpenClaw's session persistence and heartbeat system aims to eliminate the cold-start problem entirely.
Safety: backup procedures vs exec approvals
This parallel is looser. My SAFETY_AND_BACKUP_PROCEDURES.md established a three-layer protection system for data: automated daily backups, pre-experiment chaos backups, and real-time safety monitoring. It included mandatory pre-operation checks before schema changes or database modifications.
OpenClaw's exec approval system addresses a different surface of the same underlying concern. It uses deny-by-default execution with allowlist patterns: every shell command the agent wants to run must be explicitly approved or match a glob pattern you've pre-authorized. Both systems recognize that AI collaborators need guardrails — mine focused on data integrity, theirs on command execution. Different threat models, same conclusion.
The blueprint that burned
On June 8, 2025, I wrote a research document titled "Persistent QRY AI Assistant — Methodology as Integrated Intelligence." It opened with a question: "What if I could skip the context because the local AI assistant lives in the computer anyway?"
The document described a three-layer architecture: environmental awareness (file system monitoring, application integration, time pattern recognition), methodology embodiment (adapting to individual workflow patterns, anti-fragile operation), and collaborative intelligence (background enhancement, predictive assistance, learning amplification). It outlined a morning workflow where the AI already knows yesterday's progress and today's priorities. It specified local-first implementation with no cloud dependencies, user agency over AI behavior, and tool integration across the stack.
This isn't retrospective framing. The uroboro captures from that day tell the story in real time:
"Beginning qryai prototype development — systematic stack exploration for persistent QRY AI assistant. Building qryai BY improving QRY tools it will use. Recursive methodology development through tool enhancement." — June 8, 13:33
"INTEGRATION SUCCESS: Phase 1 qryai ↔ wherewasi complete! Built standardized tool communication using wherewasi's tool_messages table. qryai now sends methodology insights to wherewasi for context preservation." — June 8, 15:07
"TRINITY INTEGRATION COMPLETE! qryai → wherewasi → uroboro pipeline fully operational. Scout/Scribe/Scholar architecture working perfectly." — June 8, 15:47
In a single day, I went from design document to working prototype. The "trinity" was a three-tool pipeline: wherewasi as scout (context gathering), uroboro as scribe (insight capture), and a third tool for knowledge analysis. Three specialized tools communicating through shared SQLite databases, each handling one aspect of the persistence problem.
Compare this to OpenClaw's architecture: a gateway process managing sessions (the scout), memory tools for persistent capture (the scribe), and skills for specialized knowledge work (the scholar). The same three-role pattern, implemented as a product instead of a pipeline of CLIs.
I also designed a semantic search system in the same period — local vector embeddings over project files, uroboro captures, and documentation, using Ollama for on-device embeddings and Chroma for vector storage. The interface was qry search "trading bot game" returning ranked results with similarity scores. OpenClaw shipped the same concept as memory_search with support for multiple embedding backends.
There was also a smart model routing system — a Go binary that analyzed prompts for urgency, content type, and complexity, then routed to the appropriate local model. Quick summaries went to orca-mini:3b (three-second response). Code tasks went to codellama:7b. Quality content went to llama2:13b (twenty-five seconds, better output). Cost tracking showed $5-13/month for local inference versus $60-160 for cloud APIs. OpenClaw's model profiles — routing between Claude, GPT-4, and others based on task type — is the same architectural instinct at a different resource tier.
Five days later, on June 13, I made a strategic pivot: instead of building qryai as a separate tool, absorb its AI capabilities directly into uroboro. The "god-wyrm" vision — uroboro consuming the toolset into a plugin architecture — follows the same logic that led OpenClaw to build everything into a single gateway process. Consolidation is the natural endpoint when you have too many small tools solving adjacent problems.
But the same systematic testing that built the toolset also exposed its limits. When I subjected qryai's "methodology transfer" system to honest evaluation, the numbers that had looked like exponential growth — 47, 63, 95, 137 transfers — turned out to be a self-ingestion loop. The algorithm was reprocessing its own outputs each session, generating inflated counts. The actual transfers were templated suggestions: "Implement graceful degradation patterns" applied identically to every Go project, regardless of context. Confidence scores of 0.17. Not intelligence — structured automation wearing an intelligence costume.
I wrote a case study about it at the time, titled "A Reality Check." The conclusion: "useful systematic reminders packaged as AI insights." The value was real — context preservation, silent workflow integration, local-first privacy — but the revolutionary AI claims weren't. The fix was unglamorous: limit scope to one day, add deduplication, throttle to three insights per session. The system worked after that. It just wasn't what I'd told myself it was.
Then the maps burned. Following the engineering wisdom of knowing when to stop exploring and start building, I archived the research documents, the prototype code, and the strategic plans into a directory called qry-archive-2025-06. The archive README quotes "The Codeless Code" about burning maps after successfully crossing the wasteland. The vision was internalized. The specific implementation was shelved.
Eight months later, OpenClaw shipped the product version of what I'd prototyped as a pipeline of markdown files and Go CLIs. The architecture was the same. The scale was different.
Where the approaches diverge
The parallels are interesting, but the divergences are more revealing. They expose different values producing different designs from the same problem constraints.
Precision vs volume
Uroboro asks you to be deliberate. When you capture a decision, you're making a choice about what's worth remembering. The system enhances your input — auto-detecting the project from git, suggesting tags from content analysis — but the human initiates the capture. This is intentional: the North Star document explicitly lists AI coaching, productivity dashboards, and gamification as anti-features. The philosophy is that you should decide what matters.
OpenClaw's memory system optimizes for availability. Embed everything, search semantically later. Multiple embedding backends (OpenAI, Gemini, Voyage), configurable result limits, citation modes. It's designed for scale and recall — you shouldn't have to worry about whether you captured the right thing, because the system will find relevant context when you need it.
Both work. They serve different relationships with information. One trusts human judgment about importance; the other trusts search infrastructure. Who decides what's worth remembering is still an open question.
Philosophy as driver vs philosophy as feature
The QRY toolset has deep philosophical roots. The ten principles in QRY_NORTHSTARS draw on Ivan Illich's concept of convivial tools, permacomputing's sustainability ethos, and a specific stance on institutional trauma as a source of systematic thinking. The code exists to embody these ideas. Local-first isn't a feature — it's an ethical position about data sovereignty and the right to own your own tools.
OpenClaw is philosophically pragmatic. Self-hosting is a feature: "your hardware, your rules." It's good engineering and it respects users, but it's downstream of product decisions rather than upstream of them. The SOUL.md template is thoughtful and humane, but it's pragmatic rather than ideological.
This maps to a real tension in tool design. Philosophy-driven tools have coherence but limited reach. Product-driven tools have reach but can lose coherence. They optimize for different outcomes.
Individual craft vs product at scale
The QRY infrastructure was built by one person for one person. The procedures are human-maintained markdown because there's one human maintaining them. Uroboro's core simplification — cutting from 10,372 lines to 6,362, removing analytics and AI bloat — was an act of editorial discipline. Features were removed because they obscured the core mission of making your work visible.
OpenClaw is built for adoption. Gateway architecture, multi-channel routing (WhatsApp, Telegram, Slack, Discord, Signal, iMessage, Matrix), plugin systems, remote macOS node support, ClawHub as a skills marketplace. It's designed for many users with many configurations. A different design space entirely.
The core architectural patterns are the same despite this difference in scope. Identity documents, temporal context, persistent memory, session lifecycle, safety boundaries. Scale changes the implementation but not the underlying problem set.
Resource consciousness
Uroboro runs entirely on local compute. SQLite, Go binary, no external API calls unless you choose to use Claude Code with the MCP integration. The permacomputing influence shows: the tool should work on constrained hardware, offline, indefinitely.
OpenClaw relies on cloud LLMs (Claude Pro/Max with Opus is recommended), multiple embedding providers for semantic search, and a gateway server that's always running. It's more capable, but it's also more hungry. The resource profile reflects different assumptions about what infrastructure you can take for granted.
Same pressure, same shape
When two independent builders arrive at the same architecture from different starting points, the architecture is probably responding to real structural constraints rather than fashion. The constraints here are well-known to anyone using LLMs as daily tools:
- LLMs are stateless — they need external memory
- LLMs have no identity — they need persistent orientation documents
- LLMs have no temporal awareness — they need time injection
- LLMs are passive — they need proactive scheduling
- LLMs are powerful — they need safety boundaries
Anyone who uses these tools seriously as a daily driver hits these walls in roughly the same order. The solutions are almost inevitable once you've identified the problems clearly enough. That's optimistic: it means the personal AI assistant pattern is real, not hype. The architecture is emerging from genuine constraints, not from investor decks.
The "You Could've Invented OpenClaw" gist makes this explicit. It walks through building the architecture from first principles, and at each step the design decision feels obvious. That's the hallmark of good engineering: when you explain it, the audience thinks they could have done it themselves. Some of us did, in our own way, with our own constraints.
The tool that forgot itself
While writing this article, I did the obvious thing: asked uroboro about its own history. Searched for captures about its development, decisions made during the build, blockers hit in those first weeks of June 2025. A tool designed to make development work visible, queried about the work that made it.
Nothing. The current uroboro database has twenty decisions, all from recent work on this site — font choices, template patterns, article formatting. The origin story was gone from the active database. A git commit from June 8 says "76 captures preserved" — important enough to build a backup script for — but when I searched for them through the tool's normal interface, they didn't exist.
Then I found the SQLite files.
Buried in an excavated archive directory, uroboro's backup databases were sitting exactly where the backup script had put them. Not 76 captures — 586, spanning June 5 through June 18, 2025. The tool's own development story, captured by the tool, preserved by a backup system the tool motivated me to build. Some of what I found:
"it works! i can tell because i forgot the project flag, cool!" — June 5, testing the new SQLite backend
"CONTEXT LOSS INCEPTION: While building miqro (context loss prevention tool), I just demonstrated perfect context loss by putting the voice_prompt.sh file in wrong directory" — June 6, experiencing the problem while building the solution
"THE BIG QUESTION: Can uroboro consume itself? Train on my capture patterns, speak as I speak, predict when I would capture?" — June 6, asking what OpenClaw's self-extensible skills system would later attempt to answer
"Sjiek represents the genesis moment — original eureka project that established QRY Labs systematic methodology" — June 7, naming the progression I'm describing in this article
The data wasn't lost. It was just invisible to the tool's current interface, sitting in backup files that the tool's own discipline had created. The active database had been reset somewhere in the iterations that followed. The archival layer survived.
This distinction — between active memory and archival memory — turns out to be the hard problem. OpenClaw has it too: JSONL sessions get compacted and pruned, daily memory files accumulate, and the system relies on semantic search to surface what matters from the pile. Claude Code compresses conversation history as it approaches context limits. Every system that stores context eventually has to decide what to keep in working memory versus what to archive, and those decisions are where the unsolved problem lives. The capture mechanism works. The retrieval mechanism works. The curation mechanism — deciding what deserves to stay in active memory across months and years, not just sessions — that's the part nobody has figured out.
What's still missing
A few hard problems remain unsolved in both approaches. These will define the next generation of this kind of tooling.
The memory quality problem. Who decides what's worth remembering? Uroboro puts that on the human, which doesn't scale. OpenClaw's semantic search casts a wide net, which creates noise. I found out what happens when you try to automate this: qryai's methodology transfer system, left unchecked, created a self-ingestion loop — the algorithm reprocessing its own outputs, inflating metrics exponentially while the actual suggestions became more generic with each cycle. The fix required scope limits, deduplication, and throttling just to stop the system from poisoning itself with its own exhaust. Any memory system that accumulates without curation will eventually drown in its own recall.
The trust spectrum. When should an agent act autonomously versus ask for permission? OpenClaw's exec approval system with allowlists is the most sophisticated answer I've seen, but it's still fundamentally a binary (allowed or not) applied to a continuous problem. Trust should probably be earned incrementally, adjusted per-domain, and revocable. We're not there yet.
The second-pass problem. Both approaches assume the human catches errors and blind spots. But the value of a collaborator is partly in seeing what you can't. Neither system has a good mechanism for the AI to push back on the human's assumptions or flag patterns the human might be missing — not from a safety perspective, but from a craft perspective. There's a deeper issue here: AI collaboration is fundamentally reductive. I documented this in June 2025 — sessions that start with detailed, context-aware analysis and end with boilerplate and recursive suggestion loops. Without fresh creative input from a human who's actively learning, the collaboration degrades toward noise. The implication for both architectures: persistent memory and session management are necessary but not sufficient. The system also needs a human who keeps growing, or the AI's output converges on the mean.
Resource consciousness. OpenClaw is resource-intensive by design. My approach was frugal by philosophy. As these tools become daily infrastructure, the energy and cost profiles matter. Permacomputing and personal AI are going to have to find common ground, and it's not obvious how.
Building small tools to understand problems deeply has value even if — especially if — the market builds the product version later. The understanding doesn't go to waste. It makes you a better user of whatever comes next, because you know which problems are structural and which are implementation choices. You know where the bodies are buried because you buried some of them yourself.
The fact that a solo developer with markdown files and a Go CLI arrived at the same architectural patterns as a well-resourced open source project doesn't mean I predicted the future. It means the problems were sitting there in plain sight for anyone paying attention. The tools were obvious. The execution was the hard part, and OpenClaw executed at a scale I never attempted.
But I know these problems from the inside now. And that knowledge compounds.