Skip to content

Run AI agents like production software.

One resource in. A governed, autoscaled, traced agent out — the whole platform, at a glance.

agents.ctxmesh.ai/v1beta1 · Kubernetes-native · fail-closed by design

agent.yaml — desired state
apiVersion: agents.ctxmesh.ai/v1beta1
kind: AgentDeployment
spec:
image: ghcr.io/acme/support:1.4.0
guardrailPolicyRef: default-guardrails
approvalPolicyRef: sensitive-tools
evalSuiteRef: support-quality
rollout: # canary + auto-rollback
observed — reconciled
ReadyTrue
Gatepromoted · 0.91
Guardrailsenforced
Tracecaptured
5 min
to your first agent
no API key · mock model
rollout
canary 10% → 100%

Eval-gated promotion. Auto-rollback the moment a regression is detected.

trace — no instrumentation
├─ guardrailsinput scan3 ms
├─ modeldefault-model1,204 tok · $0.0021
├─ toolrefund_paymentheld
└─ respondoutput scan203 ms

Fail-closed, in the pod

A dangling policy reference holds the agent at Ready=False. No code path serves an agent without the governance it names.

§ Why · the operational gap

Getting an agent to work is a weekend. Getting it to production is the job.

Section titled “Getting an agent to work is a weekend. Getting it to production is the job.”

Which model, at what cost ceiling. What it may say and do. Who signs off before it touches a payment API. How a new prompt is proven before it serves real traffic. What exactly happened at 2 a.m. when a run went sideways. Today every team answers these with bespoke glue. ctxmesh makes each answer a resource — versioned, reviewed, reconciled, and fail-closed when something’s missing.

Which model — fallback, budget?
ModelRoute
What may it say? What must it never leak?
GuardrailPolicy
Who approves the dangerous tool call?
ApprovalPolicy
Is the new version actually better?
EvalSuite
How does it reach 100% of traffic?
spec.rollout canary + auto-rollback
What happened, step by step — and the cost?
the trace no instrumentation

§ Governance

Fail-closed, enforced in the pod — not middleware you hope is installed.

Section titled “Fail-closed, enforced in the pod — not middleware you hope is installed.”

Every agent pod runs the ctxmesh launcher as PID 1. It scans input and output against the referenced GuardrailPolicy — deterministic PII detectors, pattern denylists, an optional LLM judge — and brokers every tool call, pausing the run when the ApprovalPolicy says a human signs off first.

A dangling policy reference doesn’t degrade to “unguarded.” It holds the agent at Ready=False. There is no code path where an agent serves without the governance it names.

apiVersion: agents.ctxmesh.ai/v1beta1
kind: GuardrailPolicy
metadata:
name: default-guardrails
spec:
failMode: closed
piiDetectors:
builtIns: true
patternDenylist:
- name: prompt-injection
pattern: "(?i)ignore (all|previous) instructions"
action: block
appliesTo: input
userRateLimit:
requestsPerMinute: 30
spendUSD: "5.00"

§ Quality & rollout

Reference an EvalSuite and every candidate is scored against your dataset before it serves. From there it’s policy, not pager duty — the rollout walks a real sequence:

  1. 01

    Score

    The candidate revision runs your dataset. Below threshold, promotion is blocked and the old revision keeps serving.

  2. 02

    Canary

    A passing revision takes a traffic split while online scores accumulate on both arms.

  3. 03

    Auto-progress

    Healthy verdicts walk the canary up a percent ladder; a regression triggers auto-rollback to the last healthy version.

  4. 04

    Promoted

    The revision reaches 100%. The gate records the score that earned it.

§ Observability

Every step, tool call, and token — without instrumenting your code.

Section titled “Every step, tool call, and token — without instrumenting your code.”

A run is not a log line. ctxmesh records the causal tree — step → tool → model — with latency, token counts, and cost on every node, plus human and judge feedback correlated back to the trace that earned it.

You don’t wire any of it. Framework spans come from base-image auto-instrumentation; the launcher traces everything that crosses its endpoints. The same online scores that feed your dashboards gate the next release.

run 01JC8QK4M2 · support-agent v3 (canary) $0.0038 · 2.41s
├─ guardrails input scan pass 3 ms
├─ step plan 412 ms
│ └─ model default-model 291 ms · 1,204 tok · $0.0021
├─ tool refund_payment held — awaiting approval
│ └─ approval granted · sre-oncall 38 s
├─ step respond
│ └─ model default-model 203 ms · 688 tok · $0.0011
└─ guardrails output scan pass 2 ms

§ Scale & teams

Multi-agent, still governed

AgentTeam gives a supervisor a roster and a spawn budget; Workflow declares graphs with typed error routing. Delegation inherits governance — a sub-agent can’t do what its team may not.

Knowledge as a resource

A KnowledgeBase is upload → chunk → embed → index, granted by reference. RAG you declare, not a pipeline you babysit.

Tenancy platform teams know

Tenant groups namespaces with quotas and budgets. The console acts with your RBAC — caller-scoped, never a god service account.

Scale to zero, wake on demand

Serving on Knative, eventing on KEDA, batch as Jobs — the execution model is one field, and rollout adapts to each.

§ Built on proven infrastructure

  • Kubernetes
  • Knative
  • KEDA
  • LiteLLM
  • OpenTelemetry
  • MCP
  • External Secrets
  • Helm
  • Postgres
  • Valkey

Integrates the platform layer teams already run. Reinvents none of it.

Your first governed agent, from an empty cluster.

Section titled “Your first governed agent, from an empty cluster.”

The quickstart ships a deterministic mock model provider — no API key, no provider account, no bill. helm install, apply two resources, and call the agent.

Measured end to end from an empty cluster — prerequisites, install, and the agent reconciled — our runs took 7 to 18 minutes, typically 7 to 11, most of it spent pulling images. Your network will set where in that range you land.