Diagrid Catalyst¶
Diagrid Catalyst is an enterprise-grade agentic durable execution, security, and governance platform purpose-built for deploying resilient AI agents and multi-agent workflows at scale with native FastMCP 3.1 protocol support.
What it is¶
Diagrid Catalyst is a serverless, durable execution platform designed to run AI workloads, agents, and long-running workflows with built-in checkpointing, self-recovery, WebAssembly (Wasm) isolated micro-runtimes, and zero-trust security. Built on the open-source Distributed Application Runtime (Dapr) and its high-performance workflow engine, Catalyst intercepts agent executions and tool calls in real time. It enables stateful, autonomous systems to seamlessly survive crashes, cloud redeployments, and API outages without repeating expensive model invocations or losing intermediate execution state.
What problem it solves¶
AI agents often execute complex, multi-step chains of thought, tool invocations, and API calls. In high-concurrency or long-running tasks, single-point failures—such as transient network errors, rate limits, server redeployments, or container crashes—traditionally cause the entire agent loop to fail. If an agent crashes at step 99 of a 100-step task, restarting it from scratch wastes massive latencies, token consumption, and API costs.
Diagrid Catalyst solves this by: - Zero-loss Failures: Saving intermediate agent execution states, FastMCP 3.1 tool inputs, and model outputs at every step. - Self-Healing Loops: Intercepting agent runners to replay from the last completed check-point or tool execution when restarted. - Security & Governance Gateways: Providing mTLS, zero-trust cryptographic identities, and fine-grained Model Context Protocol (FastMCP 3.1) policy control on every agent and downstream tool call. - Action Attestation: Creating a tamper-proof cryptographic audit trail to prove what an autonomous agent actually did.
Where it fits in the stack¶
Infrastructure / Durable Execution Layer. Diagrid Catalyst sits as an orchestration, security, and persistence wrapper around popular agent frameworks (like LangGraph, CrewAI, Microsoft Agent Framework, or FastMCP 3.1 multi-agent topologies) and downstream services (databases, custom tools, or cloud resources).
┌─────────────────────────────────────────────────────────┐
│ User / Multi-Agent Applications │
│ (FastMCP 3.1, LangGraph, CrewAI, Google ADK) │
└────────────────────────────┬────────────────────────────┘
│ Intercepts Loop Cycles
┌────────────────────────────▼────────────────────────────┐
│ DIAGRID CATALYST │
│ (Durable Execution, mTLS, Cryptographic Auditing) │
└────────────────────────────┬────────────────────────────┘
│ Secure Tool Calls (FastMCP 3.1)
┌────────────────────────────▼────────────────────────────┐
│ Downstream Tools / VectorDBs / Databases / Cloud Infra │
└─────────────────────────────────────────────────────────┘
Model & Framework Compatibility (Early 2027 SOTA)¶
| Feature / Metric | Diagrid Catalyst 2.4 | Temporal Cloud Agentic | Restate Agent Runner | Prefect 3.0 Agentic |
|---|---|---|---|---|
| FastMCP 3.1 Support | Native real-time streaming wrapper | External adapter required | Custom plugin required | Custom middleware |
| Durable Checkpointing | Automatic Dapr state store | Code-first activity replay | Journaling KV state | Task cache state |
| Micro-Runtime Isolation | Wasm micro-containers & containers | Container / VM basis | Light event handlers | Container / Process |
| Cryptographic Audit | Native non-repudiable audit logs | Client-side tracing | Standard OTel tracing | Standard OTel tracing |
| Resume Latency | < 15ms | ~150ms | ~40ms | ~300ms |
Typical use cases¶
- Long-Horizon Multi-Agent Tasks: Executing multi-hour software engineering, research, or data aggregation agents where transient failures are highly probable.
- Enterprise Agent Security: Enforcing zero-trust network boundaries and mTLS encryption on agents interacting with critical company databases or intranet systems.
- Cryptographic Compliance and Auditing: Generating secure, non-repudiable logs of agent actions for regulatory compliance in finance, healthcare, or security sectors.
- Cost-Optimized Agent Orchestration: Preventing duplicate model and tool invocations across system restarts.
Strengths¶
- FastMCP 3.1 Integration: Seamlessly intercepts and durably records FastMCP 3.1 tool invocations across distributed microservices.
- Multi-Framework Support: Seamlessly integrates with FastMCP 3.1, LangGraph, CrewAI, Google ADK, Microsoft Agent Framework, and custom Python loops.
- Fine-Grained Checkpoint Playback: Leverages Dapr Workflows under the hood to replay orchestrations while instantly resolving previously completed activities.
- Enterprise-Grade Identity: Cryptographic identity verification and attestation at each tool and agent boundary.
Limitations¶
- Ecosystem Constraints: Requires integrating with supported agent runners, which may introduce minor overhead in simple single-step scripts.
- Side Effect Handling: External side effects must be designed carefully; downstream tools must support idempotency or reconciliation.
When to use it¶
- When your agents execute high-cost, multi-step operations (e.g., autonomous software engineering, complex database migrations).
- When agents require strict security controls, audit logs, and secure access to databases/tools.
- When deploying agents to Kubernetes/production environments where system crashes and auto-scaling events are common.
When not to use it¶
- For lightweight, single-step LLM questions or basic chatbot interfaces.
- For local-only, strictly air-gapped home environments with zero cloud connectivity or Dapr support.
Getting started¶
1. Installation¶
# Add the Diagrid Catalyst Helm repository
helm repo add diagrid https://charts.diagrid.io
helm repo update
# Install Diagrid Catalyst in your Kubernetes lab/production cluster
helm install catalyst diagrid/catalyst \
--namespace diagrid-system \
--create-namespace \
--set joinToken="YOUR_DIAGRID_CLOUD_JOIN_TOKEN"
2. Verify Deployment¶
kubectl get pods -n diagrid-system
CLI examples¶
Diagrid provides a CLI for managing, inspecting, and tracking durable agent sessions.
Managing Agent Sessions¶
# List all active durable agent sessions
diagrid sessions list
# Describe a specific failed agent session to pinpoint the exact failure step
diagrid sessions describe agent-session-091a4
# Force resume a suspended agent workflow from its last saved activity
diagrid sessions resume agent-session-091a4
Securely Invoking Tools (FastMCP 3.1)¶
# Register a secure FastMCP 3.1 tool endpoint with policy controls
diagrid tools register --name "customer-db" --url "http://mcp-server.internal:5005" --protocol fastmcp-v3.1 --policy zero-trust
API examples¶
Programmatic Durable State Validation (Python & Pydantic v2)¶
This Python script showcases how an agentic application can validate and register task execution states, checkpoint payloads, and recovery metadata using Pydantic v2 prior to handing off execution loops to Diagrid Catalyst runners.
import uuid
from datetime import datetime, timezone
from typing import Dict, Any, List, Optional
from pydantic import BaseModel, Field, field_validator, ConfigDict
class ToolExecutionRecord(BaseModel):
model_config = ConfigDict(extra="forbid", frozen=True)
tool_name: str = Field(..., description="The name of the invoked tool")
arguments: Dict[str, Any] = Field(default_factory=dict, description="Input parameters passed to the tool")
result_hash: str = Field(..., description="Cryptographic hash of the tool's return payload")
execution_time: datetime = Field(default_factory=lambda: datetime.now(timezone.utc), description="Timestamp of invocation")
class AgentCheckpointState(BaseModel):
model_config = ConfigDict(extra="forbid", frozen=True)
session_id: str = Field(..., description="Unique UUID for the durable execution session")
current_step: int = Field(..., ge=0, description="The sequential index of the active step")
completed_tools: List[ToolExecutionRecord] = Field(default_factory=list, description="List of successfully completed tools")
framework: str = Field(..., description="The underlying agent framework (e.g. FASTMCP, LANGGRAPH, CREWAI)")
state_variables: Dict[str, Any] = Field(default_factory=dict, description="Agent's current memory dictionary")
@field_validator("framework")
@classmethod
def validate_framework_choice(cls, value: str) -> str:
allowed = {"FASTMCP", "LANGGRAPH", "CREWAI", "GOOGLE_ADK", "OPENAI_AGENTS", "MICROSOFT_AGENT_FRAMEWORK", "CUSTOM"}
if value.upper() not in allowed:
raise ValueError(f"Framework must be one of {allowed}")
return value.upper()
def create_durable_checkpoint(session_id: str, step: int, tool_runs: List[dict], framework: str, state: dict) -> AgentCheckpointState:
records = []
for run in tool_runs:
records.append(ToolExecutionRecord(
tool_name=run["tool"],
arguments=run.get("args", {}),
result_hash=run["hash"]
))
checkpoint = AgentCheckpointState(
session_id=session_id,
current_step=step,
completed_tools=records,
framework=framework,
state_variables=state
)
return checkpoint
if __name__ == "__main__":
session_id = str(uuid.uuid4())
simulated_tool_runs = [
{"tool": "fetch-user-profile", "args": {"user_id": 105}, "hash": "0x91a4b83ef29"},
{"tool": "query-vector-db", "args": {"query": "durable execution"}, "hash": "0x3c7db8281fe"}
]
validated_checkpoint = create_durable_checkpoint(
session_id=session_id,
step=2,
tool_runs=simulated_tool_runs,
framework="FASTMCP",
state={"active_query": "durable execution", "user_authenticated": True}
)
print("Durable checkpoint state successfully validated for Diagrid Catalyst:")
print(validated_checkpoint.model_dump_json(indent=2))
Related tools / concepts¶
- Dapr — The Distributed Application Runtime on which Catalyst is built.
- Temporal — Code-first durable execution orchestrator.
- Model Context Protocol (MCP) — Standardized tool and resource governance protocol.
- OpenTelemetry Collector — Unified telemetry collector for Catalyst metrics.
Sources / references¶
Contribution Metadata¶
- Last reviewed: 2027-01-07
- Confidence: high