My anti-lock-in AI stack

How I separated models, harnesses, workflow and context so I can change AI providers without rebuilding my setup. Updated as the stack changes.

Published

Claude Code in the terminal was my daily driver when GLM 5.2 and Kimi K3 came out with a lot of hype around them. I had some FOMO and wanted to try them, and that's when I realized how locked in I was. My CLAUDE.md files, plugins, slash commands and settings only worked in Claude Code, so switching meant leaving my whole setup behind.

I hit the same problem during outages. When Claude went down, I could open another harness, but it had none of my skills, slash commands or context.

I solved both by rebuilding my coding workflow around OpenCode. Now switching models is a config edit. Hermes Agent runs alongside OpenCode as an always-on assistant in Slack and Telegram, and the same skills and context work in both. Both harnesses can use Claude, GPT, Gemini and several open-weight models, most of them through flat subscriptions instead of per-token billing. I rent the models and own everything else.

I keep this post current as the stack changes, and the changelog at the bottom lists what moved and when.

Open weights made lock-in expensive

Once open-weight models became good enough for daily work, staying locked into one harness meant passing them up. GLM, DeepSeek and MiMo also give me a floor, so if every closed vendor raised prices or cut off third-party tools on the same day, I'd lose some quality but keep working.

Four layers, kept apart

Lock-in enters the stack at four layers, which are the models, the harness, the workflow and the context. Vendor tools often bundle them, and adopting one product gradually tied all four of mine to it. I now keep the layers separate so I can replace one without rebuilding the others.

Models

OpenCode calls every model through a standard API endpoint, and most of those endpoints are local proxies that let it use a subscription. Meridian provides the endpoint for my Claude subscription, and CLIProxyAPI does the same for ChatGPT and my $20 Google AI Pro plan.

RouteModelsPaid with
MeridianClaude Opus, Sonnet, Fable, HaikuClaude Max subscription
CLIProxyAPIGPT-6, GeminiChatGPT subscription, Google AI Pro
OpenCode GoGLM, DeepSeek, MiMoOpenCode Go subscription
Vercel AI GatewayAnything elsePer token

GPT-6 Sol is my default today. When a new release changes the balance of quality and price, I change a config value. Each agent is defined in a markdown file with a model line that says which model it runs on. This is the line in my code reviewer's config.

model: jev-router/gpt-6-sol

Moving the reviewer to Claude means changing that line to meridian/claude-opus-5-5, and its prompt, permissions and tools stay as they are.

Harnesses

A model-independent API is not enough when the harness picks the model. SpaceXAI's Grok Bot and Hermes both give you a team of always-on agents you can reach from your phone. Grok Bot runs on a Cursor account, and its settings docs say, "Cursor manages model selection, so there's no model picker." Any routines and memory I built there would stay tied to whatever models Cursor chooses, and I couldn't swap in Opus or Sol when one of them suits a task better.

Hermes leaves the model to me. My instance currently runs GPT-6 Sol through the same ChatGPT subscription, and switching it is another config edit. It shares tools and skills with OpenCode through sync-hermes, a script in my config repo that converts OpenCode's MCP servers and skill directory into Hermes's format. A git hook reruns the script whenever the config changes, so a skill I add to OpenCode also works in Hermes.

Workflow and context

Swapping a model or harness only works if the workflow and context come along. I store skills in SKILL.md files and instructions in AGENTS.md files, both plain markdown formats that most coding agents already read. My personal context, including the style guide agents use when writing as me, lives in a git repo with a search index. Past agent sessions are indexed too, so an agent can look up a decision from last week. None of this depends on a vendor's memory feature.

Subscriptions are cheap and fragile

Flat subscriptions are the cheapest way I've found to run this much agent work. They're also the least dependable part of the stack, because Meridian and CLIProxyAPI route consumer subscriptions into tools the vendors didn't build. The vendors set the terms, and they can change them or cut off that access.

I can take the discount because no workflow depends on a particular subscription. If one is cut off or goes down, I give the affected agents a new model value, and their skills and context stay as they are.

What's still locked in

Three dependencies remain.

  • Jev Router, which selects the reasoning effort for each GPT call, depends on a GPT-6 API feature, so its savings are specific to OpenAI.
  • Meridian depends on the Claude Agent SDK, so if Anthropic changes its authentication, my Claude usage goes back to per-token pricing.
  • Most of my daily work currently runs on GPT. I can move it within minutes, but an OpenAI outage would stop that work until I do.

Changelog

  • 2026-09-28. First version. Dropped the Z.AI Coding Plan, since OpenCode Go covers GLM.
All writingEmail me