v1 Runtime infrastructure for AI agents

Control the loop. Ship the agent.

Agent RT is a lightweight, provider-neutral runtime for Python and TypeScript. Tools, streaming, policy, memory, sandboxing, orchestration, and observability stay explicit—so your agent remains understandable when the prototype becomes a system.

pip install agent-rt
agent_rt · runtime.trace live
run_42 agent=assistant tools=2 budget=max_turns 8
  1. 01text_delta"Checking the SLA first…"12ms
  2. 02tool_call_deltasearch_docs(query="SLA")schema ✓
  3. 03policy.checkside_effect=readallow
  4. 04tool_completedsearch_docs → 3 results38ms
  5. 05checkpointturn_02 committedsaved
  6. 06approval.requestsend_email → humanpaused
  7. ✓termination_reasonwaiting_for_approval
turns
1,284 tokens 6 spans
Python3.10—3.14
Node.js20 · 22 · 24 · 26
LicenseMIT
TransportWebSocket + HTTP/SSE
  • Tool calling
  • Streaming
  • Policy & approvals
  • Memory
  • Checkpoints
  • Deterministic replay
  • Sandboxed execution
  • MCP + connectors
  • Vector DB retrieval
  • Multi-agent orchestration
  • Tracing
  • Token & cost ledgers
  • OpenAI-compatible API
01 Runtime

Agent code is easy.
Runtime semantics are not.

A production agent is more than a prompt and an API call. It has to decide what can run, preserve state, recover from failure, stream useful progress, keep side effects bounded, and explain what happened afterward. Agent RT turns those concerns into explicit contracts instead of hidden framework behavior.

01

Explicit control

Termination, retries, deadlines, budgets, routing, and tool visibility are runtime decisions you can inspect and test.

02

Security boundaries

Tools, secrets, sandboxes, approvals, and data egress sit behind policy—not model-generated intent.

03

Portable primitives

Python and TypeScript share the same conceptual runtime while provider-specific behavior stays at the edge.

02 Capabilities

A small core with
serious control surfaces.

Use the pieces you need. Keep optional integrations optional.

TRANSPORT

Streaming that stays in the loop

Provider-neutral text, reasoning, status, and tool-call deltas flow through the same multi-turn runtime semantics.

Persistent Responses WebSocket fast path · HTTP/SSE fallback
POLICY

Tools are security boundaries

Typed schemas, deny-overrides permissions, model guardrails, approval gates, rate limits, deadlines, and side-effect classes execute outside the model.

STATE

Resume without guessing

Versioned workflow state, session memory, durable checkpoints, event streams, idempotency, and deterministic replay.

EXECUTION

Sandbox by contract

Command and code execution runs behind explicit resource and network policies with fail-closed backend selection.

OBSERVABILITY

Trace what the agent actually did

Hierarchical traces, privacy-aware logs, runtime metrics, token and cost ledgers, plus replay bundles for reproducible debugging.

ORCHESTRATION

Compose agents without losing the plot

Agents-as-tools, isolated subagents, supervisor/worker teams, routers, swarms, map/reduce, and speculative branches share bounded execution contracts.

OPENAI-COMPATIBLE API

Expose your agent through the API clients already speak

The optional Python server package maps configured agents to OpenAI model IDs while keeping Agent RT tools, guardrails, routing, and run limits in control.

pip install "agent-rt[api]"
GET/v1/modelsmodels
POST/v1/chat/completionsSSE
POST/v1/responsesSSE · WS
optional bearer auth · FastAPI / Uvicorn
03 Architecture

Provider-neutral inside.
Specific at the edge.

Keep application logic stable while models, tools, storage, and deployment backends evolve independently.

Your app
Agent configuration Business logic User experience
Agent RT
AgentLoop ToolRegistry Context Policy Memory Events Replay Evaluation
Adapters
Model providers MCP + connectors Sandbox backends Vector DB + storage + telemetry
stable runtime contract replaceable integration
04 Benchmarks

Measure framework overhead.
Keep the model out of it.

Agent RT ships reproducible localhost benchmarks so runtime overhead can be inspected separately from provider and network latency.

PYTHON · 30 RUNS HTTP/SSE BASELINE
1.56ms

Warm completion

Median framework/transport overhead against a deterministic localhost OpenAI-compatible mock.

TYPESCRIPT · 30 RUNS HTTP/SSE BASELINE
1.85ms

Warm completion

Median framework/transport overhead against the same class of deterministic local mock.

FEATURE BENCH · 30 RUNS PYTHON / TYPESCRIPT
3.34msPython
2.70msTypeScript

Tool round trip

One two-request tool cycle with serialized JSON tool output and common wire-level work.

CROSS-FRAMEWORK COMPARISON

Same mock. Same workload. Alternatives included.

Agent RT Alternative lower is better
PYTHON Runtime + feature + install footprint
runtime/features: 30 runs · install: 2026-10-03

Warm completion

retained HTTP/SSE runtime baseline
Agent RT 1.56 ms
LangChain 9.03 ms
LlamaIndex 4.26 ms

Tool round trip

HTTP/SSE feature benchmark
Agent RT 3.34 ms
LangChain 9.81 ms
LlamaIndex 7.31 ms

Fresh install footprint

clean uv venv · 2026-10-03
Agent RT 46.43 MiB
LangChain 69.87 MiB
LlamaIndex 219.91 MiB
TYPESCRIPT Runtime + feature + install footprint
runtime/features: 30 runs · install: 2026-10-03

Warm completion

retained HTTP/SSE runtime baseline
Agent RT 1.85 ms
LangChain.js 2.60 ms
LlamaIndex.TS 2.23 ms

Tool round trip

HTTP/SSE feature benchmark
Agent RT 2.70 ms
LangChain.js 3.33 ms
LlamaIndex.TS 2.96 ms

Fresh install footprint

clean npm install · 2026-10-03
Agent RT 1.24 MiB
LangChain.js 82.44 MiB
LlamaIndex.TS 60.38 MiB

Bars are normalized within each metric and language; printed values use the native unit for that metric.

Methodology

Runtime and feature values are checked-in 30-run medians captured on September 28, 2026 against deterministic localhost OpenAI-compatible mocks. Warm completion comes from the retained historical HTTP/SSE runtime baseline; tool round trip comes from the HTTP/SSE feature benchmark. Fresh-install footprints were measured on October 3, 2026 and include transitive dependencies: Python uses clean Python 3.13.13 uv environments with copy mode and reports logical bytes above the empty-venv baseline; TypeScript uses clean Node.js 22.22.2/npm 10.9.7 projects and reports logical node_modules bytes, with Agent RT installed from its npm-pack tarball. Agent RT's runtime runner now defaults to the persistent Responses WebSocket path, so new WebSocket runs are not compared against that historical baseline until a new 30-run WebSocket baseline is promoted. These measurements are scenario-specific framework/adapter overhead, not model latency or an overall framework score.

05 Quickstart

One runtime.
Two languages.

Start with the provider-neutral loop, then add tools, policies, storage, sandboxes, or orchestration only when the application needs them.

Moving an existing Python app? Compatibility submodules for LangChain, LlamaIndex, OpenAI, and Anthropic preserve common chat/completion entry points—including OpenAI Responses/streaming, Anthropic Messages streaming/tool history, sandbox-backed code execution, and remote MCP via Agent RT MCP clients—while routing through Agent RT. Read the migration guide ↗

  1. 01

    Install agent-rt.

  2. 02

    Set OPENAI_MODEL and your OpenAI-compatible endpoint/key as needed.

  3. 03

    Run the agent through AgentLoop.

import asyncio
import os
from agent_rt import (
    AgentConfig, AgentLoop, ContentPart,
    ModelMessage, ModelSettings, load_model,
)

async def main():
    provider = load_model(validate=False)
    agent = AgentConfig(
        name="assistant",
        instructions="Be concise and precise.",
        model=ModelSettings(model=os.environ["OPENAI_MODEL"]),
    )
    message = ModelMessage(
        role="user",
        content=(ContentPart(type="text", text="Hello."),),
    )

    result = await AgentLoop(provider).run(agent, [message])
    print(result.termination_reason)

asyncio.run(main())
provider-neutral core MIT licensed
Ready when your prototype isn't

Make the runtime
part of the design.

Build agents with explicit state, boundaries, recovery, and observability from the first serious version.