Explicit control
Termination, retries, deadlines, budgets, routing, and tool visibility are runtime decisions you can inspect and test.
Agent RT is a lightweight, provider-neutral runtime for Python and TypeScript. Tools, streaming, policy, memory, sandboxing, orchestration, and observability stay explicit—so your agent remains understandable when the prototype becomes a system.
A production agent is more than a prompt and an API call. It has to decide what can run, preserve state, recover from failure, stream useful progress, keep side effects bounded, and explain what happened afterward. Agent RT turns those concerns into explicit contracts instead of hidden framework behavior.
Termination, retries, deadlines, budgets, routing, and tool visibility are runtime decisions you can inspect and test.
Tools, secrets, sandboxes, approvals, and data egress sit behind policy—not model-generated intent.
Python and TypeScript share the same conceptual runtime while provider-specific behavior stays at the edge.
Use the pieces you need. Keep optional integrations optional.
Provider-neutral text, reasoning, status, and tool-call deltas flow through the same multi-turn runtime semantics.
Typed schemas, deny-overrides permissions, model guardrails, approval gates, rate limits, deadlines, and side-effect classes execute outside the model.
Versioned workflow state, session memory, durable checkpoints, event streams, idempotency, and deterministic replay.
Command and code execution runs behind explicit resource and network policies with fail-closed backend selection.
Hierarchical traces, privacy-aware logs, runtime metrics, token and cost ledgers, plus replay bundles for reproducible debugging.
Agents-as-tools, isolated subagents, supervisor/worker teams, routers, swarms, map/reduce, and speculative branches share bounded execution contracts.
The optional Python server package maps configured agents to OpenAI model IDs while keeping Agent RT tools, guardrails, routing, and run limits in control.
pip install "agent-rt[api]"
/v1/modelsmodels/v1/chat/completionsSSE/v1/responsesSSE · WSKeep application logic stable while models, tools, storage, and deployment backends evolve independently.
Agent RT ships reproducible localhost benchmarks so runtime overhead can be inspected separately from provider and network latency.
Median framework/transport overhead against a deterministic localhost OpenAI-compatible mock.
Median framework/transport overhead against the same class of deterministic local mock.
One two-request tool cycle with serialized JSON tool output and common wire-level work.
Bars are normalized within each metric and language; printed values use the native unit for that metric.
Runtime and feature values are checked-in 30-run medians captured on September 28, 2026 against deterministic localhost OpenAI-compatible mocks. Warm completion comes from the retained historical HTTP/SSE runtime baseline; tool round trip comes from the HTTP/SSE feature benchmark. Fresh-install footprints were measured on October 3, 2026 and include transitive dependencies: Python uses clean Python 3.13.13 uv environments with copy mode and reports logical bytes above the empty-venv baseline; TypeScript uses clean Node.js 22.22.2/npm 10.9.7 projects and reports logical node_modules bytes, with Agent RT installed from its npm-pack tarball. Agent RT's runtime runner now defaults to the persistent Responses WebSocket path, so new WebSocket runs are not compared against that historical baseline until a new 30-run WebSocket baseline is promoted. These measurements are scenario-specific framework/adapter overhead, not model latency or an overall framework score.
Start with the provider-neutral loop, then add tools, policies, storage, sandboxes, or orchestration only when the application needs them.
Moving an existing Python app? Compatibility submodules for LangChain, LlamaIndex, OpenAI, and Anthropic preserve common chat/completion entry points—including OpenAI Responses/streaming, Anthropic Messages streaming/tool history, sandbox-backed code execution, and remote MCP via Agent RT MCP clients—while routing through Agent RT. Read the migration guide ↗
Install agent-rt.
Set OPENAI_MODEL and your OpenAI-compatible endpoint/key as needed.
Run the agent through AgentLoop.
import asyncio
import os
from agent_rt import (
AgentConfig, AgentLoop, ContentPart,
ModelMessage, ModelSettings, load_model,
)
async def main():
provider = load_model(validate=False)
agent = AgentConfig(
name="assistant",
instructions="Be concise and precise.",
model=ModelSettings(model=os.environ["OPENAI_MODEL"]),
)
message = ModelMessage(
role="user",
content=(ContentPart(type="text", text="Hello."),),
)
result = await AgentLoop(provider).run(agent, [message])
print(result.termination_reason)
asyncio.run(main())
const {
AgentLoop,
loadModel,
} = require("agent-rt");
async function main() {
const provider = await loadModel(process.env, { validate: false });
const agent = {
name: "assistant",
instructions: "Be concise and precise.",
model: { model: process.env.OPENAI_MODEL },
};
const message = {
role: "user",
content: [{ type: "text", text: "Hello." }],
};
const result =
await new AgentLoop(provider).run(agent, [message]);
console.log(result.terminationReason);
}
main();
Build agents with explicit state, boundaries, recovery, and observability from the first serious version.