Private beta

From broken
to merged.

AgentReplay turns a production failure into a tested, verified pull request. You just review and merge.

Read-only to verify · writes blocked by construction · a human merges every change
incident INC-218 → pull/412
Dashboard preview

Works with the stack you already run

OpenTelemetryLangGraphOpenAI Agents SDKCrewAIPydantic AILangSmithLangfuseVercel AI SDKraw logs OpenTelemetryLangGraphOpenAI Agents SDKCrewAIPydantic AILangSmithLangfuseVercel AI SDKraw logs

Your agent will fail in production.
The only question is what happens next.

The evidence is scattered.The failure lives across traces, logs and a Slack thread — and it never runs the same way twice.
You can't safely reproduce it.Re-running the incident means real refunds, real emails, real database writes. So you don't.
The fix has no proof.You patch a prompt, ship it, and hope. Nothing shows the bug is gone — or that it stays gone.
Dashboards only watch.Observability tells you what happened. It never fixes it, and the same class of bug returns next month.

If your agent did it,
we can replay it.

Any framework, any trace source. AgentReplay turns a real production failure into a fixed, tested, verified pull request — without ever touching production.

Detects incidents on its own

Point your OpenTelemetry exporter at us, or paste a trace. Errors, loops and unsafe writes become triaged incidents — no SDK, no code change.

Reproduces in a sealed sandbox

The incident re-runs in an isolated microVM with no credentials, so write-capable tools are blocked by physics — not by a setting someone can forget.

Regression fixtures

Every fix leaves a test in your repo. The incident that hurt you once can never quietly return.

Verified fix PRs

Not a guess — a pull request whose fix was proven by execution before it reached you.

Framework-neutral

We integrate with the trace you already emit — LangGraph, OpenAI SDK, CrewAI or custom. No lock-in.

/ 01

Detect

An agent misbehaves in production. A trace arrives — via OpenTelemetry, a pasted export, or raw logs — and becomes a triaged incident with the failing step pinned.

OTel · LangSmith · Langfuse · paste
incident/INC-218 · critical
Incident detail
/ 02

Reproduce

The failure re-runs in a sealed microVM with write-capable tools blocked. First we prove the incident is real — the regression test fails on the original code.

Fargate microVM · egress denied
sandbox/run · sealed
Phase 1 · reproduce
/ 03

Fix & verify

AgentReplay writes the code fix and a real regression test, then proves it by execution: the test that failed on the original code now passes on the patch. Verified, not guessed.

reproduced ∧ fixed = verified
fix-inc-218.diff · + test
Phase 2 · verify
/ 04

Ship

A pull request lands with the fix, the test, and the audit log — on a branch, never your main. You review it on your laptop or your phone, and merge. A human ships every change.

read-only access · human merges
pull/412 · open
Pull request

What sets us apart

Observability shows you the crash. AgentReplay ships the fix that stops it — with proof.

AgentReplay
Ships a verified fix as a pull request
Reproduces the incident in a sealed sandbox
Write-capable tools blocked by construction
Leaves a regression test in your repo
Framework-neutral — no lock-in
Observability dashboards
Show you charts; the fixing is yours
Replay is for looking, not for testing
No side-effect safety when you re-run
Nothing durable — the bug can return
Often tied to one framework
the incident that paged you last week + a verified pull request, waiting for review

Bring one incident.

The private beta starts with a single real failure from your agent. We reproduce it, fix it, prove it, and open the PR — together.

Read-only to verify Writes blocked by construction A human merges every change