# Codemode

Orchestrate tools with JavaScript, filter results, and call classifier or image models.

Codemode is an optional tool that lets the model write JavaScript to call the session's tools, run independent calls in parallel, and filter intermediate results. Only explicit script output reaches the chat model, not every nested tool result.

## Enable

```sh
zot --codemode
zot --codemode-only
zot --tools read,glob,codemode
```

`--codemode` retains direct tool definitions. `--codemode-only` hides direct client tool definitions and supplies inline TypeScript declarations for nondeferred tools instead. The callable registry and permission checks do not change. `--no-tools` takes precedence, and explicit `--tools` selection restricts the built-in tools.

Merge this into `$ZOT_HOME/config.json` for persistent configuration:

```json
{
  "codemode": {
    "enabled": true,
    "mode": "only",
    "inlineBudget": 3000
  }
}
```

`mode` accepts `on` and `only`. `inlineBudget` is a non-negative estimated token budget for inline declarations, defaulting to 3000. Zero leaves discovery available without inline declarations. CLI mode overrides the configured mode. These settings apply across interactive, print, JSON, RPC, and bot modes. Go SDK users can set `sdk.Config.Codemode` or select `"codemode"` in `Config.Tools`.

## Source and execution

Scripts are JavaScript async function bodies with top-level `await` and `return`, without Markdown fences. TypeScript declarations describe tool inputs and outputs, but scripts themselves use JavaScript. Supported GPT-5 and newer Responses endpoints advertise a native custom tool with a Lark grammar, accepting raw source. Other providers use a JSON `code` argument:

```json
{"code":"const source = await tools.read({path:'package.json'})\nreturn JSON.parse(source).name"}
```

The executor also accepts a JSON string or raw source. Native source deltas are normalized to the same JSON `code` representation for guards, events, RPC, and saved sessions. User-defined models can opt in or out with `compat.supportsOpenAIGrammarTools` in `models.json`.

Use an optional first source line to set output and timeout options:

```js
// @options: {"max_output_tokens":2000,"timeout_ms":60000}
const sources = await Promise.all([
  tools.read({path:"package.json"}),
  tools.read({path:"packages/api/package.json"})
])
return sources.map(source => JSON.parse(source).name)
```

There is no default script deadline. Caller cancellation still terminates VM execution. `timeout_ms` accepts positive integers up to 2147483647. The legacy JSON argument `timeout_ms` remains supported, with source options taking precedence. Each invocation starts a fresh VM. Ordinary variables do not survive between calls.

## Script API

| API | Behavior |
| --- | --- |
| `await tools.name(args)` | Call a tool with a JSON object and resolve its result. |
| `ALL_TOOLS` | Callable catalog, including deferred tools, as name and description entries. Descriptions include TypeScript declarations. |
| `await searchTools(query, {limit, namespace})` | BM25-ranked discovery, with a default limit of 8. Options are optional. |
| `await describeTool(name)` | TypeScript declaration and description, or `undefined` if unknown. |
| `await describeNamespace(name)` | Namespace metadata and tool names, or `undefined` if unknown. |
| `text(value)` | Emit text, JSON-serializing objects and arrays. |
| `console.log(...values)` | Emit one space-separated line. `info`, `warn`, `error`, and `debug` behave the same way. |
| `return value` | Emit a final value unless it is `undefined`. |
| `image(value)` | Emit a supported local base64 image. |
| `exit()` | Finish successfully immediately, retaining store writes. Script catch blocks cannot intercept it. |
| `store(key, value)` | Store a JSON copy under a string key. `undefined` deletes it. |
| `load(key)` | Return a fresh JSON copy, or `undefined` for a missing key. |

Tool names are normalized to ASCII JavaScript identifiers, replacing unsupported characters with underscores. Original names also work through `tools["custom-tool"`]. The first tool wins when normalized names collide. `tools`, `models`, `ALL_TOOLS`, and `console` are immutable. Missing members report suggestions. Use `"name" in tools` to check availability.

## Tool results

A tool declaring an output schema and returning structured content resolves to that JSON value, including structured failure results. Otherwise, text blocks are joined with newlines. Failed or refused calls without a structured result reject with their error text. Tools without structured output do not automatically forward images.

Shell tools return structured results:

```ts
{
  output: string
  truncated: boolean
  full_output_path?: string
  exit_code: number
  wall_time_seconds: number
}
```

`output` contains merged stdout and stderr. Beyond 1 MiB, it keeps the first and last 512 KiB around a byte-omission marker, trimmed to UTF-8 boundaries. `truncated` describes that script output, not the smaller direct-tool display. `full_output_path` is present only when script output was truncated and spilling the full stream to a private temporary file succeeded. Nonzero exits resolve so scripts can inspect `exit_code`. Cancellation and timeouts reject.

```js
const result = await tools.bash({command:"go test ./..."})
text({exit:result.exit_code, output:result.output})
```

## Images and persistent state

`image` accepts a base64 data URL, an object with `image_url: "data:..."`, or an object with `type: "image"`, `data`, and `mimeType`. PNG, JPEG, GIF, and WebP are recognized from signatures. Remote image URLs are rejected.

Store keys must be strings and values must be JSON-serializable. Each JSON value is limited to 262144 UTF-16 characters, with 1048576 characters across keys and JSON values. Limit violations throw `RangeError`. Reads and writes copy JSON, so later mutation does not change stored values.

```js
const count = load("runs") ?? 0
store("runs", count + 1)
return {runs:load("runs")}
```

Writes commit only when the script succeeds, including `exit()`. Errors and cancellation discard that script's writes. Completed external side effects are not rolled back. Snapshots persist in message metadata under `tool_state:codemode`, separate from model-visible content. Resume and branch replay reconstruct state from selected messages, and compaction carries snapshots forward. Saved transcripts may contain private application state.

## Classifier and image models

Catalog methods are `models.getModelsOfType(type, provider?)`, `models.getAvailableOfType(type, provider?)`, and `models.getModelOfType(type, provider, id)`. An unknown model returns `undefined`. Types are `chat`, `classifier`, and `image`. Availability means credentials or a no-auth provider are configured, not that an endpoint is reachable. Chat models can be listed but cannot be executed from scripts.

```js
const available = await models.getAvailableOfType("classifier")
if (available.length) {
  return await models.classify(available[0], {
    state:{text:"A synthetic example"},
    questions:{positive:{
      type:"bool",
      instructions:"Is the sentiment positive?",
      criteria:{true:"Positive sentiment", false:"Negative or neutral sentiment"}
    }}
  })
}
```

`models.classify(model, {state, questions})` requires an object state and a nonempty question map. Each question has string `instructions` and one of these types:

- `choice`: `criteria` maps labels to descriptions. Answers contain `choice`, `probabilities`, and `confidence`.
- `bool`: `criteria` contains `true` and `false` descriptions. Answers contain `probability` for true.
- `score`: `criteria` is a list of level descriptions. Answers contain `score` and `confidence`.

Supported classification transports are TypeSafe-compatible System One, Cloudflare Workers AI System One, and llama.cpp raw next-token probabilities. Cloudflare resolves `CLOUDFLARE_ACCOUNT_ID` in the host. Local choice classification supports 2 to 62 distinctly tokenized labels, and local score classification supports 2 to 10 levels.

```js
const available = await models.getAvailableOfType("image")
if (available.length) {
  const result = await models.generateImages(available[0], {
    input:[{type:"text",text:"Draw a simple landscape"}]
  })
  for (const block of result.output) {
    if (block.type === "image") image(block)
    else if (block.type === "text") text(block.text)
  }
}
```

Image generation uses OpenRouter and requires a nonempty `input` array of text or base64 image blocks. Generated images are not automatically emitted. Use `image(block)` to show them. Codemode adds a note when image generation returned images but the script emitted none.

Both inference APIs return `api`, `provider`, `model`, `timestamp`, and `stopReason`, plus `answers` or `output`. Service failures return `stopReason: "error"` and `errorMessage`, cancellation returns `"aborted"`, and invalid arguments reject. Reported usage is charged to the session even if the script later fails. Usage events carry `auxiliary: true` (Go SDK `Event.Auxiliary`) and do not replace the chat context estimate. Unknown pricing does not mean free inference. At most four inference calls run concurrently per script.

Credentials and endpoints stay in the host. Scripts select catalog models, not arbitrary endpoints or API keys. TypeSafe supports `TYPESAFE_API_KEY`. Other providers use normal credential resolution. Additional models in `models.json` use `type`, `api`, and optional `output` fields. Auxiliary API IDs are `typesafe-system-one`, `cloudflare-workers-ai-system-one`, `llama-cpp-classify`, and `openrouter-images`.

## Discovery and extensions

Extension namespaces can include descriptions and instructions for discovery ranking and `describeNamespace`. Inline declarations share a UTF-16-based budget, selecting the cheapest remaining tool from each namespace in rounds. Tools that do not fit remain discoverable.

Tool exposure defaults to `direct`. `codemode` makes a tool callable from scripts without an eager direct definition. `deferred` keeps it discoverable without inline declarations until activated. `model-only` excludes it from scripts. In `only` mode, direct client definitions are hidden, while deferred definitions remain available for later activation. See [Extensions](/docs/extensions) for structured results and registration fields.

## Output and permissions

Text defaults to 10000 estimated tokens, using four UTF-16 characters per token. `max_output_tokens` accepts a non-negative safe integer. Overflow retains the head and tail and writes full text to a private temporary file. Those files are not automatically deleted after the result is returned. Output starts with `Script completed` or `Script failed` and elapsed wall time. Failure preserves earlier emitted output.

Each invocation has separate hard limits of 16777216 UTF-16 characters across text and base64 image data and 100000 output items. Exactly reaching a limit is allowed. Exceeding either terminates the script even inside catch or finally blocks, retains earlier output, discards store writes, and cancels outstanding host calls. Increasing `max_output_tokens` does not raise these limits. Write larger data to a file with a tool instead.

Nested calls retain normal guards, argument rewrites, confirmations, permissions, and tool lifecycle events. IDs are `<parent-id>/<number>`, and parallel calls can finish out of order. Their results are observable events, not separate model transcript messages. Completed nested calls remain visible beneath their parent after completion and session reload. Arguments, status, and text results persist as display-only `nested_tool_calls` metadata, including failed scripts. Images are retained as captions. This metadata is not sent to providers, but saved and exported sessions may contain private intermediate results. Older sessions without these records cannot reconstruct nested calls.

Inference uses the same guarded execution path, with lifecycle names `models.classify` and `models.generateImages`. Catalog queries do not create tool lifecycle rows. Unawaited calls are cancelled when a script ends. Cleanup waits for host calls, so a tool ignoring cancellation can delay cleanup beyond the script deadline. Recursion is refused, with a maximum nested tool chain of four.

## Isolation and Go execution package

JavaScript runs using Goja inside an embedded WebAssembly VM executed by Wazero. No external runtime, CGO, or subprocess is required. Each invocation has fresh linear memory with a hard 256 MiB cap, including the worker's Go runtime and JavaScript heap. Oversized allocations terminate the worker rather than necessarily raising a catchable exception. Context cancellation forcibly stops VM execution.

The VM mounts no filesystem and exposes no direct network, process, module-loading, or timer APIs. Outside effects require explicit host capabilities. This confines worker execution, not authorized tools. The memory cap does not cover host-side JSON decoding, output buffering, concurrent VM totals, or tool allocations. It is not an operating-system sandbox for tools. [Jail mode](/docs/jail) remains an accident-prevention guardrail.

Go applications can use `github.com/patriceckhart/zot/packages/codemode` independently of the agent, core, and provider packages. `Initialize()` compiles the shared worker, and `Run(ctx, Start, Caller, emit)` starts a fresh VM. The caller owns host permissions and state persistence. Host callbacks may run concurrently and must honor cancellation. The agent adapter lives in `packages/agent/codemode`. For execution-package details and worker regeneration commands, see [https://github.com/patriceckhart/zot/blob/main/docs/codemode.md](https://github.com/patriceckhart/zot/blob/main/docs/codemode.md).
