Scaling agents
Goal: attach an elastic scaling rule to an agent — you declare the intent (a trigger and bounds) and the platform selects the backend.
Prerequisites: the platform installed; an agent already deployed in the
same namespace. For queue-depth, the agent should be an eventing consumer;
for schedule, an executionModel: job agent.
Two scaling surfaces
Section titled “Two scaling surfaces”There are two distinct ways to scale an agent — use the right one:
AgentDeployment.spec.scaling(inlinemin/max) — the basic Knative autoscaler bounds of a plainservingagent (request-driven, defaultsmin: 0,max: 3). Use this for a simple request/response agent that just needs bounds. See Deploy an agent.AgentScalingPolicy(this guide) — a separate resource that declares a trigger (request-rate / custom-metric / queue-depth / schedule) plusmin/max, and lets the controller pick the backend (Knative autoscaling annotations, a KEDAScaledObject, or a CronJob). Use this for event-driven, metric-driven, or scheduled scaling.
How it works
Section titled “How it works”An AgentScalingPolicy targets one AgentDeployment in the same namespace by name (spec.agentRef). The
controller reads spec.trigger and generates the matching backend:
spec.trigger |
Backend the controller generates |
|---|---|
request-rate |
Knative autoscaling annotations (concurrency/rps) on the agent’s ksvc |
custom-metric |
Knative custom autoscaling (class/metric annotations) on the ksvc |
queue-depth |
a KEDA ScaledObject targeting the agent’s deployment, reading the registry broker depth |
schedule |
a CronJob (for an executionModel: job agent) on the cron expression |
You declare intent and bounds; the platform selects the mechanism and reports which one on
status.backend.
1. Author the policy
Section titled “1. Author the policy”A queue-depth policy with scale-to-zero (min: 0):
apiVersion: agents.ctxmesh.ai/v1beta1kind: AgentScalingPolicymetadata: name: worker-queue namespace: my-teamspec: agentRef: worker-agent # the AgentDeployment (same namespace) to scale trigger: queue-depth # request-rate | custom-metric | queue-depth | schedule min: 0 # 0 = scale-to-zero when the queue is empty max: 20 cooldown: 60s # cooldown after a scale event (Go duration) # queueRef: { name: support-broker } # optional; defaults to the registry brokerApply it:
kubectl apply -f worker-queue.yaml2. Watch it attach
Section titled “2. Watch it attach”kubectl get agentscalingpolicy worker-queue -n my-team -wkubectl get agentscalingpolicy worker-queue -n my-team \ -o jsonpath='{.status.conditions[?(@.type=="Ready")].status} {.status.backend}{"\n"}'# → True keda-scaledobjectReady=True means the target agent exists and the chosen backend resource was created/updated;
status.backend reports which one (knative-annotations, keda-scaledobject, or cronjob). Under load,
replicas scale toward max; when idle, back toward min (to 0 when min: 0).
Other triggers
Section titled “Other triggers”Request-rate — for an interactive serving agent that should keep warm capacity:
apiVersion: agents.ctxmesh.ai/v1beta1kind: AgentScalingPolicymetadata: name: triage-rate namespace: my-teamspec: agentRef: triage-agent trigger: request-rate min: 1 # keep one warm replica (no cold start) max: 10Schedule — run an executionModel: job agent on a cron (the controller emits a CronJob):
apiVersion: agents.ctxmesh.ai/v1beta1kind: AgentScalingPolicymetadata: name: nightly-batch namespace: my-teamspec: agentRef: batch-agent # executionModel: job trigger: schedule max: 1 schedule: "0 2 * * *" # required when trigger: schedule (5-field cron)When to use / when not
Section titled “When to use / when not”- Use
AgentScalingPolicyfor queue-depth (event consumers), custom-metric, request-rate warm capacity, or a scheduled job. - Use the inline
AgentDeployment.spec.scalingfor the plain min/max bounds of a request-driven agent — don’t reach for a policy when you only need bounds. - Not a way to exceed org limits: the scaler bounds pods; provider rate limits stay coordinated by the model gateway, so scaling out never breaches org quotas.
Defaults
Section titled “Defaults”mindefaults to 0 (scale-to-zero);maxis required (minimum 1, and must be ≥min).cooldowndefaults to60s(Go duration^[0-9]+(s|m|h)$).scheduleis required whentrigger: schedule, ignored otherwise.queueRefdefaults to the agent’s registry broker; only meaningful forqueue-depth.- Scaling never drops below
min—min: 1guarantees a warm replica;min: 0allows idle-to-zero.
Failure modes
Section titled “Failure modes”The Ready condition carries the reason on failure:
AgentNotFound—spec.agentRefnames anAgentDeploymentthat doesn’t exist in the namespace. Deploy it, or fix the ref.InvalidTrigger— an unsupported trigger, ortrigger: schedulewithout aschedule(rejected by the admission CEL rule). Set a valid trigger / add the cron.BackendError— the controller could not create/update the chosen backend (KEDAScaledObject, Knative annotations, or CronJob).kubectl describeshows the underlying cause.