Codemode

Orchestrate tools with JavaScript, filter results, and call classifier or image models.

Codemode is an optional tool that lets the model write JavaScript to call the session's tools, run independent calls in parallel, and filter intermediate results. Only explicit script output reaches the chat model, not every nested tool result.

Enable

sh
zot --codemode
zot --codemode-only
zot --tools read,glob,codemode

--codemode retains direct tool definitions. --codemode-only hides direct client tool definitions and supplies inline TypeScript declarations for nondeferred tools instead. The callable registry and permission checks do not change. --no-tools takes precedence, and explicit --tools selection restricts the built-in tools.

Merge this into $ZOT_HOME/config.json for persistent configuration:

json
{
  "codemode": {
    "enabled": true,
    "mode": "only",
    "inlineBudget": 3000
  }
}

mode accepts on and only. inlineBudget is a non-negative estimated token budget for inline declarations, defaulting to 3000. Zero leaves discovery available without inline declarations. CLI mode overrides the configured mode. These settings apply across interactive, print, JSON, RPC, and bot modes. Go SDK users can set sdk.Config.Codemode or select "codemode" in Config.Tools.

Source and execution

Scripts are JavaScript async function bodies with top-level await and return, without Markdown fences. TypeScript declarations describe tool inputs and outputs, but scripts themselves use JavaScript. Supported GPT-5 and newer Responses endpoints advertise a native custom tool with a Lark grammar, accepting raw source. Other providers use a JSON code argument:

json
{"code":"const source = await tools.read({path:'package.json'})\nreturn JSON.parse(source).name"}

The executor also accepts a JSON string or raw source. Native source deltas are normalized to the same JSON code representation for guards, events, RPC, and saved sessions. User-defined models can opt in or out with compat.supportsOpenAIGrammarTools in models.json.

Use an optional first source line to set output and timeout options:

js
// @options: {"max_output_tokens":2000,"timeout_ms":60000}
const sources = await Promise.all([
  tools.read({path:"package.json"}),
  tools.read({path:"packages/api/package.json"})
])
return sources.map(source => JSON.parse(source).name)

There is no default script deadline. Caller cancellation still terminates VM execution. timeout_ms accepts positive integers up to 2147483647. The legacy JSON argument timeout_ms remains supported, with source options taking precedence. Each invocation starts a fresh VM. Ordinary variables do not survive between calls.

Script API

APIBehavior
[[await tools.name(args)]]Call a tool with a JSON object and resolve its result.
[[ALL_TOOLS]]Callable catalog, including deferred tools, as name and description entries. Descriptions include TypeScript declarations.
[[await searchTools(query, {limit, namespace})]]BM25-ranked discovery, with a default limit of 8. Options are optional.
[[await describeTool(name)]]TypeScript declaration and description, or undefined if unknown.
[[await describeNamespace(name)]]Namespace metadata and tool names, or undefined if unknown.
[[text(value)]]Emit text, JSON-serializing objects and arrays.
[[console.log(...values)]]Emit one space-separated line. info, warn, error, and debug behave the same way.
[[return value]]Emit a final value unless it is undefined.
[[image(value)]]Emit a supported local base64 image.
[[exit()]]Finish successfully immediately, retaining store writes. Script catch blocks cannot intercept it.
[[store(key, value)]]Store a JSON copy under a string key. undefined deletes it.
[[load(key)]]Return a fresh JSON copy, or undefined for a missing key.

Tool names are normalized to ASCII JavaScript identifiers, replacing unsupported characters with underscores. Original names also work through tools["custom-tool"]. The first tool wins when normalized names collide. tools, models, ALL_TOOLS, and console are immutable. Missing members report suggestions. Use "name" in tools to check availability.

Tool results

A tool declaring an output schema and returning structured content resolves to that JSON value, including structured failure results. Otherwise, text blocks are joined with newlines. Failed or refused calls without a structured result reject with their error text. Tools without structured output do not automatically forward images.

Shell tools return structured results:

ts
{
  output: string
  truncated: boolean
  full_output_path?: string
  exit_code: number
  wall_time_seconds: number
}

output contains merged stdout and stderr. Beyond 1 MiB, it keeps the first and last 512 KiB around a byte-omission marker, trimmed to UTF-8 boundaries. truncated describes that script output, not the smaller direct-tool display. full_output_path is present only when script output was truncated and spilling the full stream to a private temporary file succeeded. Nonzero exits resolve so scripts can inspect exit_code. Cancellation and timeouts reject.

js
const result = await tools.bash({command:"go test ./..."})
text({exit:result.exit_code, output:result.output})

Images and persistent state

image accepts a base64 data URL, an object with image_url: "data:...", or an object with type: "image", data, and mimeType. PNG, JPEG, GIF, and WebP are recognized from signatures. Remote image URLs are rejected.

Store keys must be strings and values must be JSON-serializable. Each JSON value is limited to 262144 UTF-16 characters, with 1048576 characters across keys and JSON values. Limit violations throw RangeError. Reads and writes copy JSON, so later mutation does not change stored values.

js
const count = load("runs") ?? 0
store("runs", count + 1)
return {runs:load("runs")}

Writes commit only when the script succeeds, including exit(). Errors and cancellation discard that script's writes. Completed external side effects are not rolled back. Snapshots persist in message metadata under tool_state:codemode, separate from model-visible content. Resume and branch replay reconstruct state from selected messages, and compaction carries snapshots forward. Saved transcripts may contain private application state.

Classifier and image models

Catalog methods are models.getModelsOfType(type, provider?), models.getAvailableOfType(type, provider?), and models.getModelOfType(type, provider, id). An unknown model returns undefined. Types are chat, classifier, and image. Availability means credentials or a no-auth provider are configured, not that an endpoint is reachable. Chat models can be listed but cannot be executed from scripts.

js
const available = await models.getAvailableOfType("classifier")
if (available.length) {
  return await models.classify(available[0], {
    state:{text:"A synthetic example"},
    questions:{positive:{
      type:"bool",
      instructions:"Is the sentiment positive?",
      criteria:{true:"Positive sentiment", false:"Negative or neutral sentiment"}
    }}
  })
}

models.classify(model, {state, questions}) requires an object state and a nonempty question map. Each question has string instructions and one of these types:

  • choice: criteria maps labels to descriptions. Answers contain choice, probabilities, and confidence.
  • bool: criteria contains true and false descriptions. Answers contain probability for true.
  • score: criteria is a list of level descriptions. Answers contain score and confidence.

Supported classification transports are TypeSafe-compatible System One, Cloudflare Workers AI System One, and llama.cpp raw next-token probabilities. Cloudflare resolves CLOUDFLARE_ACCOUNT_ID in the host. Local choice classification supports 2 to 62 distinctly tokenized labels, and local score classification supports 2 to 10 levels.

js
const available = await models.getAvailableOfType("image")
if (available.length) {
  const result = await models.generateImages(available[0], {
    input:[{type:"text",text:"Draw a simple landscape"}]
  })
  for (const block of result.output) {
    if (block.type === "image") image(block)
    else if (block.type === "text") text(block.text)
  }
}

Image generation uses OpenRouter and requires a nonempty input array of text or base64 image blocks. Generated images are not automatically emitted. Use image(block) to show them. Codemode adds a note when image generation returned images but the script emitted none.

Both inference APIs return api, provider, model, timestamp, and stopReason, plus answers or output. Service failures return stopReason: "error" and errorMessage, cancellation returns "aborted", and invalid arguments reject. Reported usage is charged to the session even if the script later fails. Usage events carry auxiliary: true (Go SDK Event.Auxiliary) and do not replace the chat context estimate. Unknown pricing does not mean free inference. At most four inference calls run concurrently per script.

Credentials and endpoints stay in the host. Scripts select catalog models, not arbitrary endpoints or API keys. TypeSafe supports TYPESAFE_API_KEY. Other providers use normal credential resolution. Additional models in models.json use type, api, and optional output fields. Auxiliary API IDs are typesafe-system-one, cloudflare-workers-ai-system-one, llama-cpp-classify, and openrouter-images.

Discovery and extensions

Extension namespaces can include descriptions and instructions for discovery ranking and describeNamespace. Inline declarations share a UTF-16-based budget, selecting the cheapest remaining tool from each namespace in rounds. Tools that do not fit remain discoverable.

Tool exposure defaults to direct. codemode makes a tool callable from scripts without an eager direct definition. deferred keeps it discoverable without inline declarations until activated. model-only excludes it from scripts. In only mode, direct client definitions are hidden, while deferred definitions remain available for later activation. See Extensions for structured results and registration fields.

Output and permissions

Text defaults to 10000 estimated tokens, using four UTF-16 characters per token. max_output_tokens accepts a non-negative safe integer. Overflow retains the head and tail and writes full text to a private temporary file. Those files are not automatically deleted after the result is returned. Output starts with Script completed or Script failed and elapsed wall time. Failure preserves earlier emitted output.

Each invocation has separate hard limits of 16777216 UTF-16 characters across text and base64 image data and 100000 output items. Exactly reaching a limit is allowed. Exceeding either terminates the script even inside catch or finally blocks, retains earlier output, discards store writes, and cancels outstanding host calls. Increasing max_output_tokens does not raise these limits. Write larger data to a file with a tool instead.

Nested calls retain normal guards, argument rewrites, confirmations, permissions, and tool lifecycle events. IDs are <parent-id>/<number>, and parallel calls can finish out of order. Their results are observable events, not separate model transcript messages. Completed nested calls remain visible beneath their parent after completion and session reload. Arguments, status, and text results persist as display-only nested_tool_calls metadata, including failed scripts. Images are retained as captions. This metadata is not sent to providers, but saved and exported sessions may contain private intermediate results. Older sessions without these records cannot reconstruct nested calls.

Inference uses the same guarded execution path, with lifecycle names models.classify and models.generateImages. Catalog queries do not create tool lifecycle rows. Unawaited calls are cancelled when a script ends. Cleanup waits for host calls, so a tool ignoring cancellation can delay cleanup beyond the script deadline. Recursion is refused, with a maximum nested tool chain of four.

Isolation and Go execution package

JavaScript runs using Goja inside an embedded WebAssembly VM executed by Wazero. No external runtime, CGO, or subprocess is required. Each invocation has fresh linear memory with a hard 256 MiB cap, including the worker's Go runtime and JavaScript heap. Oversized allocations terminate the worker rather than necessarily raising a catchable exception. Context cancellation forcibly stops VM execution.

The VM mounts no filesystem and exposes no direct network, process, module-loading, or timer APIs. Outside effects require explicit host capabilities. This confines worker execution, not authorized tools. The memory cap does not cover host-side JSON decoding, output buffering, concurrent VM totals, or tool allocations. It is not an operating-system sandbox for tools. Jail mode remains an accident-prevention guardrail.

Go applications can use github.com/patriceckhart/zot/packages/codemode independently of the agent, core, and provider packages. Initialize() compiles the shared worker, and Run(ctx, Start, Caller, emit) starts a fresh VM. The caller owns host permissions and state persistence. Host callbacks may run concurrently and must honor cancellation. The agent adapter lives in packages/agent/codemode. For execution-package details and worker regeneration commands, see https://github.com/patriceckhart/zot/blob/main/docs/codemode.md.