Skip to content

Deploy an agent

Goal: get your own agent image running as a governed, autoscaled service.

Prerequisites: the platform installed; a namespace you can write to; a model available via a ModelRoute; your agent packaged on a supported base image.

An agent is one resource. The minimum is an image and an execution model:

apiVersion: agents.ctxmesh.ai/v1beta1
kind: AgentDeployment
metadata:
name: support-agent
namespace: my-team
spec:
image: ghcr.io/my-org/support-agent:1.0.0
executionModel: serving # serving | eventing | job
# The model is chosen in your code by calling MODEL_GATEWAY_URL with model="<ModelRoute name>".
scaling:
min: 0 # scale to zero when idle (default)
max: 3

Apply it:

Terminal window
kubectl apply -f support-agent.yaml
Terminal window
kubectl get agentdeployment support-agent -n my-team -w
# Ready transitions to True once the serving revision is up.
kubectl get agentdeployment support-agent -n my-team \
-o jsonpath='{.status.conditions[?(@.type=="Ready")].status} {.status.url}{"\n"}'

The controller injects the gateway URL + launcher config, wires tracing, and reports readiness and the serving URL on status.

Send a request to status.url (or use the console Playground). Every turn is traced — open the agent’s runs in the console to see the step → tool → model tree, cost, and any guardrail decisions.

Governance is opt-in and reusable. Author the policies once (see the linked guides) and reference them:

spec:
image: ghcr.io/my-org/support-agent:1.0.0
executionModel: serving
guardrailPolicyRef: default-guardrails # content rules — /guides/guardrails/
approvalPolicyRef: sensitive-tools # human approval — /guides/approvals/
evalSuiteRef: support-quality # gate releases — /guides/evals-and-the-deploy-gate/
feedbackStoreRef: support-feedback # feedback model — /guides/feedback-and-improvement/
sessionMemory:
scope: session # per-conversation memory — /guides/memory-and-sessions/

A dangling reference fails closed: the agent goes Ready=False and is held rather than served ungoverned.

  • serving for interactive request/response agents; eventing for async/event-driven work; job for batch. See Execution models.
  • Leave scaling.min: 0 for scale-to-zero unless you need warm capacity (cold-start latency vs cost).
  • executionModel defaults to serving; scaling defaults to min: 0, max: 3.
  • No evalSuiteRef ⇒ no deploy gate (zero overhead). No policy refs ⇒ today’s ungoverned behavior.
  • Image pull / crash → the revision is not Ready; kubectl describe + the pod logs show why.
  • Dangling policy ref → Ready=False with the reason on the condition (fix or remove the ref).
  • Model alias unresolved (no matching ModelRoute) → model calls fail fast at the gateway.

Connect a model provider · Guardrails · Gate a release on an eval suite · AgentDeployment reference