Memory & sessions
Goal: give an agent conversation memory so a multi-turn chat keeps its context — even across a scale-to-zero cold start.
Prerequisites: an agent deployed (Deploy an agent). For a shared
scratchpad the agent must be a member of an AgentRegistry.
How session memory works
Section titled “How session memory works”Agent state never lives in the pod — that is what makes scale-to-zero safe. Conversation context
lives in a dedicated Valkey (the State Layer), fronted by a control-plane proxy, and is reached by
the agent through the launcher’s traced localhost memory endpoint. You turn it on with one field on the
agent; the controller injects the backend wiring, and the stock managed loop replays recent turns
before each new turn and appends the completed exchange after it. Every memory op is a memory.* span
in the run’s trace tree. See Memory & state for the model.
1. Enable session memory on the agent
Section titled “1. Enable session memory on the agent”Add spec.sessionMemory to the AgentDeployment. The presence of the field is the switch — absent
means no conversation memory.
apiVersion: agents.ctxmesh.ai/v1beta1kind: AgentDeploymentmetadata: name: support-agent namespace: my-teamspec: image: ghcr.io/my-org/support-agent:1.4.0 executionModel: serving sessionMemory: scope: session # session (private, default) | shared (registry team scratchpad) # backend: # addr: ctxmesh-statelayer.ctxmesh.svc:6379 # default if omittedApply it:
kubectl apply -f support-agent.yamlEnabling (or disabling) memory changes the pod template, so the controller rolls a new Knative
revision. Ready transitions back to True once the new revision is up:
kubectl get agentdeployment support-agent -n my-team \ -o jsonpath='{.status.conditions[?(@.type=="Ready")].status}{"\n"}'# → True2. Correlate turns with a conversationId
Section titled “2. Correlate turns with a conversationId”Memory is keyed per conversation. The console chat sends one stable conversationId per session (as the
X-Conversation-Id header); the BFF forwards it, and the managed loop threads memory automatically —
only when the agent is bound to memory. Turn 2 in the same conversation sees turn 1’s context; a new
conversationId starts a fresh context.
sessionscope keys memory asmem:{namespace}/{agent}:{conversationId}— private to this agent.sharedscope keys it asmem:shared:{registry}:{conversationId}— one context that every agent in the same registry conversation reads and writes. Each writer’s messages are stamped server-authoritatively with that agent’s own name, so “who wrote what” stays recoverable in the shared log.sharedrequires registry membership; ascope: sharedon a non-member agent silently keeps the private layout (a visible misconfiguration, never a rootless shared key).
3. Verify it survives scale-to-zero
Section titled “3. Verify it survives scale-to-zero”The proof that state is external: converse, let the agent scale to zero, converse again on a cold pod, and the earlier context is still there.
# After a first turn, watch the agent scale to zero, then send a second turn on the same conversationId.kubectl get agentdeployment support-agent -n my-team -w# A cold pod's first memory GET returns the stored context — no lost history.Open the run in the console and confirm the memory.get / memory.append spans in the trace tree.
Per-user isolation (opt-in)
Section titled “Per-user isolation (opt-in)”spec.sessionMemory.perUser: true isolates each invoking end-user’s conversation into its own bucket,
so two users on the same conversationId never share history. It is product-grade (launcher-stamped
from the verified run capability), private scope only, and defaults off.
spec: sessionMemory: scope: session perUser: trueWhen to use / when not
Section titled “When to use / when not”- Use
sessionfor any multi-turn agent that should remember earlier turns in the same conversation. - Use
sharedfor a team of registry agents collaborating on one thread (a shared scratchpad). - Not for cross-conversation facts an agent should recall by meaning — that is
spec.longTermMemory(semantic, pgvector); see Memory & state. - Not for a searchable document corpus — that is a KnowledgeBase.
Defaults
Section titled “Defaults”scopedefaults tosession(private per-agent);perUserdefaults tofalse.backend.addrdefaults toctxmesh-statelayer.ctxmesh.svc:6379.- The managed loop replays only the last
MAX_HISTORY_MESSAGES= 40 turns as prompt context; older turns fall out of the prompt (the store still holds them, capped below). - The store retains the last 500 entries per conversation (
LTRIMon append) and each conversation expires after a 24-hour TTL, refreshed on every write (not on reads). Only the clean{user, assistant}pair is persisted — intermediate tool-call scratchpad messages are not.
Failure modes
Section titled “Failure modes”- Backend unreachable → the launcher’s memory endpoint returns a typed error; the agent treats memory as best-effort and answers without context rather than failing the turn (the same non-fatal philosophy as tool calls). Note: in a proxy-cutover, network-isolated install, memory fails open (silent loss) rather than blocking the turn.
- No
conversationId(or nospec.sessionMemory) → the agent runs single-shot: nothing is replayed and nothing is stored (today’s Playground behaviour). scope: sharedon a non-registry agent → no shared key is created; the agent silently keeps its private layout. Add the agent to a registry to get the shared scratchpad.- Bad
conversationId(empty,>128chars, or containing/,:, whitespace, or control characters) → the memory op is rejected (a key/path-injection guard).
See also
Section titled “See also”- AgentDeployment reference (
spec.sessionMemory,spec.longTermMemory) - Memory & state · Custom resources
- Deploy an agent · Knowledge & RAG