Runtime guide
Detailed usage notes for providers, the Python API server and CLI, sandboxes, skills, and migration entry points. The package READMEs stay short and link here.
Providers
load_model() (Python) and loadModel() (TypeScript) select a provider from the environment. OPENAI_BASE_URL selects OpenAI and takes precedence over ANTHROPIC_BASE_URL; with neither, OPENAI_MODEL alone selects OpenAI. MODEL_PROVIDER=openai|anthropic overrides auto-detection and invalid values fail fast. API keys alone, or ANTHROPIC_MODEL alone, never choose a provider.
The OpenAI provider uses the native OpenAI-compatible transport with HTTP/SSE and Responses WebSocket support; the Anthropic provider uses the optional official SDK. Reasoning controls are normalized into the provider-neutral ReasoningConfig.
Content and provider behaviour
- Media: image, PDF/file, and (OpenAI) audio parts are sent as native provider blocks;
datamay be an http(s) URL, adata:URL, base64, or bytes. Parts a provider cannot carry raise instead of being dropped. - Anthropic thinking: thinking blocks are preserved and replayed before
tool_use. Requests default tomax_tokens=4096; setmax_output_tokens(maxOutputTokensin TypeScript) for longer outputs. - Context trust: untrusted retrieved or file content reaches the model as user-role data; only workflow state and explicitly trusted items are system messages.
- Validation: tool arguments and structured output are checked against a JSON Schema subset (types, bounds,
pattern,const,anyOf/oneOf/allOf/not, local$ref,additionalProperties). - Truncation: a Responses WebSocket
response.incompletereturnsfinish_reason="length"with usage instead of raising. - Loop guarantees: every tool call is answered. When a run stops early, unexecuted calls receive a
tool_not_executedmessage and completed results from a concurrent batch are kept.
API server (Python)
Install agent-rt[api] to expose configured agents through OpenAI-compatible Chat Completions/Responses APIs and an Anthropic-compatible Messages API.
from agent_rt import AgentConfig, AgentLoop, ModelSettings, load_model
from agent_rt_api import serve
provider = load_model()
loop = AgentLoop(provider)
agent = AgentConfig(
name="assistant",
instructions="You are a helpful assistant.",
model=ModelSettings(model="your-provider-model"),
)
serve(loop, {"assistant": agent}, port=8000)The server has bounded concurrency and queues, request and generation limits, authentication, streaming, and conservative defaults for unauthenticated serving.
- Rate limiting is keyed by verified credential or client address, never by unverified header values.
- Validation happens before streaming starts; malformed bodies, bad parameters, unknown models, and unsupported content parts return a 4xx.
- Limit-driven stops report
finish_reason: "length"(Anthropicstop_reason: "max_tokens") and never expose unexecuted server-side tool calls. /healthreturns only{"status": "ok"}to unauthenticated callers when API keys are configured.
TypeScript provides the API-façade contracts but no ready-made HTTP server.
Terminal CLI (Python)
Install agent-rt[cli] to add the agent-rt command.
agent-rt
agent-rt -p "explain this stack trace"
agent-rt --session refactor --prompt "plan the migration"
agent-rt --session refactor --resumeThe CLI shares the runtime, provider routing, guardrails, tool execution, and run limits of programmatic Agent RT. Workspace writes require approval and Git integration is inspection-only: writes to .git/, .gitattributes, and .gitmodules are refused, and Git tools fail closed when repository config defines filter.* drivers. Recoverable tool errors go back to the model so the turn and its history survive.
Sandboxes
SandboxSession runs commands through a fail-closed backend selected with AGENT_RT_SANDBOX_BACKEND: native, docker, e2b, microsandbox, or swe-rex.
- Native: runs as a restricted account's primary gid unless
AGENT_RT_SANDBOX_GIDis set and never keeps the caller's gid. The TypeScript backend usessetpriv --clear-groups --no-new-privs. - Docker: all capabilities dropped,
no-new-privileges, and swap capped to the memory limit. SetAGENT_RT_SANDBOX_DOCKER_USERfor a non-root user andAGENT_RT_SANDBOX_DOCKER_READ_ONLY=truefor a read-only root filesystem. - Lifetime: timeouts and cancellation kill the process group (native) or container (Docker).
- Limits: output is capped while read (default 16 MiB) and zero memory or process limits are rejected.
- Cleanup and snapshots: call
await session.close()to release cloud sandboxes. Snapshots capture only the in-process workspace and session settings, not backend disks.
Skills and registration safety
SkillRegistry loads Claude Code and Codex style SKILL.md directories and collections. Python exposes install_from_path and install_directory.
import { SkillRegistry } from "agent-rt/extensions";
const skills = new SkillRegistry();
await skills.installFromPath(".claude/skills/review", { activate: true });
await skills.installDirectory(".agents/skills");Registration is fail-closed: deterministic scanning always runs and externally sourced registrations require a configured Decision model before the registry is mutated.
Migration entry points
Prepend the package name to a supported upstream import path. See the migration guide for the supported subset and semantic gaps.
| Upstream | Python | TypeScript |
|---|---|---|
openai | agent_rt.openai | agent-rt/openai |
anthropic / @anthropic-ai/sdk | agent_rt.anthropic | agent-rt/@anthropic-ai/sdk |
langchain_openai / @langchain/openai | agent_rt.langchain_openai | agent-rt/@langchain/openai |
llama_index.llms.openai / @llamaindex/openai | agent_rt.llama_index.llms.openai | agent-rt/@llamaindex/openai |
agents / @openai/agents | agent_rt.agents | agent-rt/@openai/agents |
autogen_agentchat.agents | agent_rt.autogen_agentchat.agents | Python only |
crewai | agent_rt.crewai | Python only |