Armin Ronacher explains Codemode: LLM tool calls written as JavaScript in a no-network QuickJS sandbox on the Pi harness
- Codemode lets the model issue and compose tool calls from a programming language (JavaScript in Pi) instead of JSON tool schemas, and it runs on the harness (brain) side inside QuickJS on a WASM runtime with no network, no file system, no timers, and limited RAM.
- Pi 1.0 added MCP support through Codemode; it is enabled by default only when MCP is on, and a user can turn it on with
"defaultTools": ["+codemode"]. The name Codemode was coined by Cloudflare. - A regular bash tool call puts only the trailing 2000 lines into context, while a Codemode invocation gets larger outputs sent structurally, and Codemode can stash state into the transcript for later calls in the session.
- Ronacher splits the system into a trusted harness (brain) and a separate execution environment (hands) that runs bash, so the two sit on different file systems with different trust levels; a sandbox like Gondolin covers the bash side but not the harness.
- Pi exposes APIs that make no sense as regular tools, such as image generation and one-shot text classification, through Codemode rather than tools that would waste context.
Hacker News opinions
Not sure why almost every codemode implementation picks JavaScript. I prototyped an agent using bash as the codemode language and it worked just as well, with literally zero prompt to teach it.
Because the models are trained on JavaScript, you get away with way fewer instructions. Promises give you concurrency for free. And code mode runs on the harness side, where bash is a tricky target.
Sandboxing bash is way harder than JS or Lua, which have fantastic embedded tooling. Every model is fully trained on JS anyway, so that needs no teaching either.
The thing I still can't figure out is recovery after interrupted execution, and how you tell completed side effects apart from calls that are safe to repeat.
How is this different from Claude writing mini scripts all the time? Claude's scripts can't call MCP tools, so everything has to be CLIs or libraries, and you lose the self-documenting input-output schema codemode is going for.
I tried codemode in a prototype and stripped it out again. My small local LLM burned more tokens, and when an embedded tool call failed on parameter validation the whole parent code block failed, so the model kept rewriting it.
That matches a fear I had about small models. It works fantastically with Opus and Sol, but I always wondered how well it scales down.
Monty from the pydantic team is a joy if you need to run unverified code securely. It's a simplified Python dialect, and you can call it from JS.
Between this and the Cloudflare post, it's a lot of words with no simple system-level picture. This is homoiconicity, actor semantics, and object capabilities rediscovered by hacking outward from token streams.
It's simpler than that. Instead of a fixed JSON schema for tool calls, let the model write a program to call and compose the tools however it wants, with APIs to reach into the harness.
Can someone tell me how to disable this permanently in Pi? I set autoEnableCodemode to false and it still won't turn off.
MCP and codemode don't auto-enable. {"extensions": ["-builtin:codemode"]} is all you need, and -builtin:mcp kills MCP independently.
Code mode seems to cause progressive lobotomisation in Sol 6.1 around subagents. The more it uses it, the worse its subagent prompts get, merging words together and spamming keywords.
I still can't parse what it actually is. The article is titled What is Codemode and then halfway through says if you are not familiar with Codemode. That passage about 2000 lines means nothing to me.
This is a very complex apparatus for little gain. Pi's promise was bash and no other tools. The minimally invasive fix would be to inject tool calls as virtual bash commands.
Am I crazy for liking codemode? It cuts the number of turns by about an order of magnitude.