Runtime
Durable agents that survive anything.
Checkpointed, resumable runs with automatic retries, branching and human-in-the-loop pauses, out of the box.
Lumen is the AI-native platform for building, tracing and deploying production agents. Go from prompt to planet-scale in seconds, not sprints.
Free forever for side projects · No credit card
Trusted by 14,000+ teams shipping agents
Running in production at 14,000+ teams, from seed-stage to Fortune 100
Platform
Runtime, memory, tools, evals and observability, designed together, so your agents are fast, safe and boring to operate.
Runtime
Checkpointed, resumable runs with automatic retries, branching and human-in-the-loop pauses, out of the box.
Inference
Requests route to the nearest warm GPU automatically.
38msp50 TTFT
312edge regions
Evals
Continuous evals gate every deploy against your real traffic.
Memory
Vector recall with TTLs, scopes and per-user isolation.
Traces
Every prompt, tool call and token, replayable step by step.
Registry
400+ typed, permissioned tools. Bring your own in one line.
Guardrails
PII redaction, injection shields and hard spend caps, enforced at the edge.
Deploy
Immutable versions, atomic traffic shifts and one-click rollbacks. No containers to babysit.
8.2s median push → live
Build
1.8s
Evals
2.4s · 248/250
Replicate
3.1s · 312 regions
Live
0.9s · traffic 100%
How it works
Point Lumen at your models, data and tools. The MCP-native registry gives every agent typed, permissioned access in a single line.
Wire steps into a durable graph. Lumen compiles it into a checkpointed runtime with retries, branching and human approval built in.
Push once. Lumen ships to 312 edge regions in under ten seconds, with traces, evals and rollbacks switched on from the very first request.
Developers
Typed SDKs for TypeScript and Python. A CLI that feels instant. A runtime that gets out of your way.
import { agent, tool, memory } from "@lumen/sdk";
import { z } from "zod";
const lookupOrder = tool({
name: "crm.lookup_order",
input: z.object({ orderId: z.string() }),
run: ({ orderId }) => crm.orders.get(orderId),
});
export default agent({
name: "support-agent",
model: "lumen/frontier-1",
memory: memory.semantic({ ttl: "30d" }),
tools: [lookupOrder, "mcp://zendesk"],
guardrails: { pii: "redact", budget: "$0.04/run" },
async run(ctx, msg) {
const plan = await ctx.plan(msg);
return ctx.stream(plan);
},
});
from lumen import agent, tool, memory
@tool("crm.lookup_order")
def lookup_order(order_id: str):
return crm.orders.get(order_id)
@agent(
model="lumen/frontier-1",
memory=memory.semantic(ttl="30d"),
tools=[lookup_order, "mcp://zendesk"],
guardrails={"pii": "redact", "budget": "$0.04/run"},
)
async def support(ctx, msg):
plan = await ctx.plan(msg)
return ctx.stream(plan)
# evals/support.yaml
suite: support-regression
agent: support-agent@latest
dataset: s3://acme/tickets/2026-q3.jsonl
scorers:
- type: llm-judge
rubric: "Resolves the issue without escalation"
- type: latency
p95: 1200ms
gate:
pass_rate: ">= 0.97"
on_fail: block-deploy
2.4B
agent runs / month
Orchestrated on Lumen across 14,000+ production teams.
38ms
median time-to-first-token
Global p50, measured from 312 edge regions.
312+
edge regions
Deploy once. Run within 40 ms of 96% of the world’s users.
99.99%
uptime, contractually
Backed by a financially guaranteed SLA on Scale.
Customers
We replaced eleven microservices and a homegrown orchestrator with Lumen in a single sprint. Our support agent now resolves 63% of tickets end-to-end, and I finally sleep through on-call.
Priya Raman
VP Engineering, Halcyon
The traces alone are worth it. I can replay any failed run, step by step, and see exactly which tool call went sideways.
Marcus Feld
Staff Engineer, Northbeam
Deploys in eight seconds. I timed it. Twice. Then I made the whole team watch.
Aiko Tanaka
Founder, Kestrel
Evals as a deploy gate changed how we ship. Regressions get caught long before a customer ever sees them.
Daniel Okoye
Head of AI, Meridian
Latency dropped from 2.1s to 380ms after we moved inference to the edge. We didn’t change a line of code.
Tom Whitaker
CTO, Parallax
Lumen is the first platform that treats agents like real distributed systems: checkpoints, retries, idempotency. It’s boring in exactly the right ways.
Sofia Lindqvist
Principal Engineer, Orbital
We evaluated five platforms. Lumen was the only one our security team approved on the first pass.
Grace Adeyemi
CISO, Quanta
Pricing
Pay for agent runs, not seats. Every plan includes traces, evals and the full runtime.
For side projects and first experiments.
$0 / month
Free forever · no card
Start for freeFor teams running agents in production.
$39
/ month
$49
Billed annually · $468/yr
Start 14-day trialFor organisations with serious traffic.
$199
/ month
$249
Billed annually · $2,388/yr
Talk to usNeed VPC deployment, data residency or a custom SLA? Talk to sales →
FAQ
Still stuck? Our engineers answer in minutes, not days.
Talk to an engineerLumen is an AI-native platform for building, running and observing production agents. It bundles a durable runtime, semantic memory, a tool registry, continuous evals and global inference behind one SDK and one deploy command.
Any of them. Lumen is model-agnostic: bring your own keys for the major providers, use open-weight models we host on our edge GPUs, or mix them per step with automatic fallbacks and cost-aware routing.
Agents compile to lightweight, snapshot-based sandboxes that are pre-warmed across 312 regions. A deploy pushes a new immutable version and shifts traffic atomically, so there are no cold builds and no container registry round-trips.
Never. Prompts, traces and memory are encrypted in transit and at rest, isolated per project, and never used to train any model. Bring-your-own-cloud and regional data residency are available on Scale.
Yes. Enterprise customers can run the data plane in their own AWS, GCP or Azure account, or on-prem, while Lumen manages the control plane. Talk to us about air-gapped deployments.
Nothing breaks. Agents keep running and overage is billed at a transparent per-run rate, with hard budget caps you control. You’ll get alerts at 70%, 90% and 100% of your quota.
Ready when you are
Start free in under a minute. No credit card, no sales call, no cold starts.