Multi-agent, still governed
AgentTeam gives a supervisor a roster and a spawn budget; Workflow declares graphs with typed error routing. Delegation inherits governance — a sub-agent can’t do what its team may not.
agents.ctxmesh.ai/v1beta1 · Kubernetes-native · fail-closed by design
apiVersion: agents.ctxmesh.ai/v1beta1 kind: AgentDeployment spec: image: ghcr.io/acme/support:1.4.0 guardrailPolicyRef: default-guardrails approvalPolicyRef: sensitive-tools evalSuiteRef: support-quality rollout: # canary + auto-rollback
Eval-gated promotion. Auto-rollback the moment a regression is detected.
A dangling policy reference holds the agent at Ready=False. No code path serves an agent without the governance it names.
§ Why · the operational gap
Which model, at what cost ceiling. What it may say and do. Who signs off before it touches a payment API. How a new prompt is proven before it serves real traffic. What exactly happened at 2 a.m. when a run went sideways. Today every team answers these with bespoke glue. ctxmesh makes each answer a resource — versioned, reviewed, reconciled, and fail-closed when something’s missing.
ModelRouteGuardrailPolicyApprovalPolicyEvalSuitespec.rollout canary + auto-rollback§ Governance
Every agent pod runs the ctxmesh launcher as PID 1. It scans input and output
against the referenced GuardrailPolicy — deterministic PII detectors, pattern
denylists, an optional LLM judge — and brokers every tool call, pausing the run when
the ApprovalPolicy says a human signs off first.
A dangling policy reference doesn’t degrade to “unguarded.” It holds the agent at
Ready=False. There is no code path where an agent
serves without the governance it names.
apiVersion: agents.ctxmesh.ai/v1beta1kind: GuardrailPolicymetadata: name: default-guardrailsspec: failMode: closed piiDetectors: builtIns: true patternDenylist: - name: prompt-injection pattern: "(?i)ignore (all|previous) instructions" action: block appliesTo: input userRateLimit: requestsPerMinute: 30 spendUSD: "5.00"§ Quality & rollout
Reference an EvalSuite and every candidate is scored against your dataset before it
serves. From there it’s policy, not pager duty — the rollout walks a real sequence:
The candidate revision runs your dataset. Below threshold, promotion is blocked and the old revision keeps serving.
A passing revision takes a traffic split while online scores accumulate on both arms.
Healthy verdicts walk the canary up a percent ladder; a regression triggers auto-rollback to the last healthy version.
The revision reaches 100%. The gate records the score that earned it.
§ Observability
A run is not a log line. ctxmesh records the causal tree — step → tool → model — with latency, token counts, and cost on every node, plus human and judge feedback correlated back to the trace that earned it.
You don’t wire any of it. Framework spans come from base-image auto-instrumentation; the launcher traces everything that crosses its endpoints. The same online scores that feed your dashboards gate the next release.
run 01JC8QK4M2 · support-agent v3 (canary) $0.0038 · 2.41s├─ guardrails input scan pass 3 ms├─ step plan 412 ms│ └─ model default-model 291 ms · 1,204 tok · $0.0021├─ tool refund_payment held — awaiting approval│ └─ approval granted · sre-oncall 38 s├─ step respond│ └─ model default-model 203 ms · 688 tok · $0.0011└─ guardrails output scan pass 2 ms§ Scale & teams
AgentTeam gives a supervisor a roster and a spawn budget; Workflow declares graphs with typed error routing. Delegation inherits governance — a sub-agent can’t do what its team may not.
A KnowledgeBase is upload → chunk → embed → index, granted by reference. RAG you declare, not a pipeline you babysit.
Tenant groups namespaces with quotas and budgets. The console acts with your RBAC — caller-scoped, never a god service account.
Serving on Knative, eventing on KEDA, batch as Jobs — the execution model is one field, and rollout adapts to each.
§ Built on proven infrastructure
Integrates the platform layer teams already run. Reinvents none of it.
The quickstart ships a deterministic mock model provider — no API key, no provider
account, no bill. helm install, apply two resources, and call the agent.
Measured end to end from an empty cluster — prerequisites, install, and the agent reconciled — our runs took 7 to 18 minutes, typically 7 to 11, most of it spent pulling images. Your network will set where in that range you land.