← Back to Explore

Getting Off the Cloud AI Treadmill

How to reduce dependency on paid AI without losing what makes it useful

The loop

Cloud AI creates a dependency loop. Productivity goes up. Bill goes up. Workflow gets tied to services you don't control. Usage limits. Price increases. Service changes. Outages during deadline week.

I hit my Claude token ceiling mid-refactor and had to stop coding. Not because the problem was hard. Because the meter ran out. That's a dependency, not a tool.

The answer isn't going fully local. Not yet — local models can't match frontier reasoning for complex architecture work. The answer is being honest about what actually needs cloud capabilities and what doesn't.

The 80/20 split

Most AI tasks don't need GPT-4. They need "good enough, fast, and available."

I run code completion locally because it needs to be instant and private. I run documentation generation locally because a 7B model handles it fine and I don't want my project notes on someone else's server. I use wherewasi for context recovery when I switch between projects — that's a local operation, no reason to phone home for it.

I send architecture questions to Claude because a local model can't hold the whole system in its head the way a frontier model can. I use cloud for novel problems that need broad knowledge — the kind of thing where you'd otherwise spend two hours reading documentation. That's maybe 20% of my actual AI interactions.

A fast local model at infinite availability beats a better model you're rationing. Especially for the stuff that's mostly pattern-matching anyway.

What local actually looks like

I built uroboro to capture development decisions as I make them. It hooks into git commits — every commit message gets tagged, categorized, and indexed in a local SQLite database. No cloud round-trip, no API cost, no privacy leak. The AI processing happens on my machine through Ollama.

Context switching used to cost me 20 minutes every time I jumped between projects. Wherewasi snapshots project state when I leave — what I was working on, what's next, what's blocking — and restores it when I come back. All local. All instant.

Sjiek automates git diff to clipboard. Tiny tool. But it eliminated the copy-paste friction between my terminal and AI conversations, which turned out to be the thing actually slowing me down.

None of this is impressive individually. The compound effect is what matters. Each tool removes a small friction, and the frictions were where all the time was going.

Hardware honesty

I'm on 8GB. It works. Not comfortable, but it works. Small quantized models run. Inference is slow. Context windows are tight. But the models are mine and they don't bill me per token.

The break-even math depends on what you're spending. If cloud AI costs you $20/month, hardware upgrades don't make economic sense. If you're burning through $200/month in API calls (I've been there), a one-time hardware investment starts looking rational pretty fast.

There's also the question nobody wants to ask: what's the environmental cost of routing every code completion through a data center? Running a 7B model on your laptop uses a fraction of the energy. That shouldn't be the only reason, but it shouldn't be zero reasons either.

What goes wrong

Local models are worse at complex reasoning. That's just true. When I need the AI to think about system design across five interacting services, local doesn't cut it. That's fine — that's why the cloud 20% exists. The mistake is trying to force local where it doesn't work, then concluding local doesn't work at all.

Setup isn't trivial. Ollama makes it manageable, but you're still configuring model selection, managing disk space, tuning context lengths. It's systems work. If you don't enjoy systems work, this path will annoy you.

And there's a real risk of over-engineering the local setup. I've caught myself spending more time optimizing the toolchain than using it. At some point you have to stop building the workshop and start building in it.

Flip from 80% cloud to 80% local. Takes months. Requires measuring what actually needs cloud and what you assumed needed cloud because that's how you started.

Not because local is ideologically pure. Because owning the infrastructure means nobody can pull the rug during deadline week.