Skip to content

Conversation overhead & provider mechanics

When an agent pane is open and you send a message, there are two distinct layers that add context to the conversation before it reaches the AI provider. Understanding both layers — and how provider-side prompt caching interacts with them — is what distinguishes “sent every turn” (technically true) from “costs full tokens every turn” (not true, after turn 1).

Layer 1 — AgentMux injects once, at agent launch

Section titled “Layer 1 — AgentMux injects once, at agent launch”

Before the provider CLI process starts, AgentMux assembles and writes a CLAUDE.md file (or equivalent) to the agent’s working directory. This file is built from:

ComponentSourcePurpose
SoulAgent definitionCore identity and behaviour instructions
Agent MDAgent definitionPer-agent task instructions
Global memory bundlesBrain bundles (ordered)Shared knowledge prepended to every agent
Skills indexInstalled skills# Available Skills section listing callable skills
Template varsRuntime{{AGENT}}, {{AGENT_DISPLAY}}, {{AGENT_ID}}, {{DATE}} substituted at write time. {{WORKING_DIR}} is a supported placeholder but is not currently wired up on this path — see the note below.

Source: agentmux-srv/src/backend/agent_config.rs build_config_files(), agentmux-srv/src/server/app_api/agent_open.rs write_agent_config_files() (single call site: the agent.open RPC handler — fires once per agent-pane open, confirmed no other caller exists).

This file is written once per agent launch, not once per turn. AgentMux does not call the provider’s API directly — it hands the file off to the CLI and steps back.

Startup payload (separate from the CLAUDE.md file): On the first turn of a new session (not on --resume), AgentMux also sends a structured user-turn message containing runtime identity, assigned accounts, peer agents, date/version, and any custom startup instructions. This is a one-time cost, not repeated. Source: frontend/app/view/agent/startup/buildStartupPayload.ts, invoked from frontend/app/view/agent/agent-view.tsx (guarded so it’s skipped whenever the block already carries a session ID, i.e. on resume).

Layer 2 — The provider CLI injects on every API call

Section titled “Layer 2 — The provider CLI injects on every API call”

The provider CLI reads its project instructions file from disk — CLAUDE.md for Claude Code, AGENTS.md for Codex, GEMINI.md for Gemini — and includes its contents as the system prompt on every HTTP request it makes to the provider’s API. From the API’s perspective, the system prompt is re-sent on every turn.

This is where prompt caching matters.

Each provider CLI manages system prompt caching using its own provider-specific mechanism. The details below cover Claude Code, which is the most documented case. Other providers use different caching APIs with different semantics.

The Claude Code CLI applies cache_control: {"type": "ephemeral"} to the system prompt block — this is Anthropic’s explicit prompt caching header (anthropic-beta: prompt-caching-2024-07-31). The Anthropic API caches the system prompt server-side after the first request; subsequent turns in the same session pay cache-read rates.

TurnToken type chargedNotes
Turn 1cache_creation_input_tokensFull cost to write the cache entry
Turn 2+cache_read_input_tokens~10× cheaper than full input tokens

AgentMux tracks all three token types — input_tokens, cache_creation_input_tokens, and cache_read_input_tokens — and sums them to show real context size. The correct formula is:

real_context_tokens = input_tokens + cache_creation_input_tokens + cache_read_input_tokens

Source: frontend/app/view/agent/providers/claude-translator.ts (token accumulation, lines 95-98), docs/specs/SPEC_CONTEXT_VISIBILITY_2026_06_17.md. The Rust-side struct is TokenCounts in agentmux-srv/src/agents/types.rs — it uses shortened field names (input, output, cache_creation, cache_read); the _input_tokens-suffixed wire names above only exist in the frontend’s raw usage parsing, not as Rust struct fields.

Note: these three fields are parsed per-turn for the live UI counters only — they are not persisted anywhere queryable. There is currently no way to check historically whether caching landed for a past session; only live, in the moment.

OpenAI applies prompt caching automatically for inputs over 1,024 tokens — no explicit header is required. Common prompt prefixes are cached server-side; the cost reduction is reflected in the cached_tokens field of the usage response. AgentMux does not currently surface Codex cache token counts separately.

Gemini supports context caching via a separate explicit API call (not an in-request header). The Gemini CLI manages this internally. Token cost reporting follows Gemini’s usage response format, which differs from Anthropic’s.

What makes AgentMux’s system prompt larger than a bare CLI session

Section titled “What makes AgentMux’s system prompt larger than a bare CLI session”

A standard claude CLI session uses whatever CLAUDE.md files it finds in the project directory hierarchy. AgentMux replaces that with its assembled file, which bundles soul + memory bundles + skills index on top of the agent’s own instructions. The assembled file is typically larger than a bare project’s CLAUDE.md.

Practical implications:

  • Turn 1 is more expensive — the cache-write covers a larger payload.
  • Per-turn cost (turn 2+) is the same model — cache-read is cache-read regardless of size (though larger = more cache-read tokens).
  • Cache invalidation resets the cost to turn-1 levels — see below.

Cache invalidation: when turn-1 cost recurs

Section titled “Cache invalidation: when turn-1 cost recurs”

The provider’s cached system prompt entry is invalidated whenever the system prompt content changes. For AgentMux agents, this happens when:

  • The agent’s soul, agentmd, or memory bundle content is edited
  • A memory bundle is added to or removed from the agent
  • The skills index changes (skill installed/removed)
  • A new session is started without --resume (the session ID changes, and the CLI opens a fresh context)

On invalidation, turn 1 of the next session pays full cache_creation_input_tokens again.

Mitigation for cross-instance cache reuse: the Claude Code provider’s launch args include --exclude-dynamic-system-prompt-sections, which moves per-machine content (cwd, env info, memory paths, git status) out of the CLI’s own dynamically-generated system-prompt sections and into the first user message instead — so the same assembled CLAUDE.md produces the same cached prompt prefix across different machines/instances, rather than each one minting its own cache entry. This only applies to content the CLI itself injects dynamically (separate from the CLAUDE.md-content invalidation triggers above), and is a no-op if an agent’s provider_flags sets a custom --system-prompt (a free-form, user-configurable field), since that flag’s own documented behavior only applies to the default system prompt.

AgentMux does not accumulate or re-send conversation history. History ownership depends on the controller type:

  • persistent (Claude Code): The CLI subprocess stays alive between turns and owns the messages array for the session. AgentMux communicates via PTY stdin and reads structured events from stdout — it never reconstructs or replays history.
  • subprocess (Codex, Gemini, others): A fresh CLI process is spawned per turn. The provider’s API holds session state server-side (via a session or thread ID), so history is not re-sent by the client on each turn either — it is referenced by ID.

In both cases AgentMux itself never holds or replays the conversation history.

Context compaction (automatic truncation when the context window fills up) is handled entirely inside the CLI. AgentMux does not trigger or control it — the CLI never reports its own context-window size, so AgentMux learns/seeds it from observed usage and applies a fixed buffer to estimate when compaction is imminent.

Compaction threshold for Claude Code CLI: contextWindow - 33,000 tokens. Source: frontend/app/store/agent-pane-state/context-window.tsconst COMPACTION_BUFFER = 33_000.

ProviderAPI call makerCaching appliedHistory ownerAgentMux API calls
Claude Code CLICLI subprocessYes (prompt-caching-2024-07-31)CLI subprocessNone
Codex CLICLI subprocessYes (provider-side)CLI subprocessNone
Gemini CLICLI subprocessYes (provider-side)CLI subprocessNone

AgentMux makes zero direct HTTP calls to any AI provider API. All provider interaction flows through the CLI subprocess via PTY.