Skip to content

Execution models

An agent’s execution model answers one question: what triggers it, and how does it scale? You choose it with a single field — AgentDeployment.spec.executionModel — and the platform picks the right Kubernetes machinery underneath. There are three models, and they exist because request/response, event-driven, and batch work have genuinely different scaling shapes; forcing them onto one backend is where platforms get brittle.

Model Trigger Backend it reconciles to Scales on
serving (default) synchronous request Knative Service (ksvc) request rate / concurrency (Knative KPA)
eventing a message on the registry broker a plain Deployment + Service + Trigger queue depth (KEDA)
job a schedule, or one-shot Kubernetes Job / CronJob not applicable (runs to completion)

serving is the default, so a plain agent needs nothing extra. The other two are opt-in for workloads that aren’t request-shaped.

Use it for anything a caller waits on: a chat agent, an HTTP API, a tool another agent calls synchronously. It reconciles to a Knative Service, which gives you request-driven autoscaling including scale-to-zero — an idle agent costs nothing, and a cold request scales it back up. Because the caller is waiting, request rate (not backlog) is the right scaling signal, and Knative’s KPA is the right autoscaler.

Use it for work that arrives as events rather than blocking calls: a consumer draining a topic, a reaction to another agent’s output, fan-out background processing. An eventing agent must be a member of an AgentRegistry; the registry owns a Knative Eventing broker, and the agent gets a Trigger subscribing it to that broker.

Here the model deliberately diverges from serving: an eventing agent is a plain Deployment, not a ksvc. The reason is a real, load-bearing trade-off discovered against a live cluster — KEDA scales on queue depth by naming the target Deployment, but a Knative ksvc hides its Deployment behind a generated revision name, so KEDA can’t find it, and even if it could, KEDA and Knative’s autoscaler would fight over the replica count. Event consumers aren’t request-driven, so Knative Serving buys them nothing. A plain Deployment is cleanly KEDA-scalable — so that’s what an eventing agent gets.

Async delivery has explicit semantics: at-least-once delivery, launcher-side idempotency (the consumer dedupes on the envelope’s messageId), best-effort ordering per conversationId, a per-registry dead-letter queue after N retries, and automatic blob offload to the object store for payloads over 256 KB (the envelope carries a reference; the launcher rehydrates it before your agent sees it). See Async & eventing.

Use it for offline or periodic work: a nightly summarizer, a batch scorer, a one-shot data pass. It reconciles to a Kubernetes Job (one-shot) or a CronJob (when a schedule-triggered AgentScalingPolicy targets it). The launcher still fronts the container — same runtime contract — and the pod exits when the agent completes. Overlapping runs are forbidden by default so a slow run never stacks on itself.

You don’t hand-wire an HPA or a ScaledObject. You declare intent with an AgentScalingPolicy — a trigger (request-rate / custom-metric / queue-depth / schedule) plus min/max bounds — and the platform selects the backend: Knative autoscaling annotations for request-rate, a KEDA ScaledObject for queue-depth, a CronJob for a schedule. min: 0 means scale-to-zero. Crucially, provider rate limits stay enforced at the Model Gateway, independent of pod scaling — the scaler bounds pods, the gateway bounds calls, so scaling out can never breach an organization’s provider limits.

A new AgentVersion rolls out differently depending on the model, because “shift traffic gradually” means different things for each:

  • serving — a traffic-split canary: the new revision takes a small percentage of live requests, both arms are scored per-version, and a human (or, opt-in, auto-progression) promotes to 100% or aborts to 0%. See Canary & rollout.
  • eventing — a shadow consumer group approach rather than a traffic split (there’s no request stream to split).
  • job — version-pinned scoring of a sample before the new version becomes the one the schedule runs.

In every case the eval gate can block a version that regresses, and the model determines only how the shift is performed.

Reach for serving unless you have a reason not to — it’s the default and covers most agents. Choose eventing when work arrives as events and you want backlog-driven scaling and a DLQ. Choose job for batch or scheduled runs that start, do finite work, and exit.