Observability

Every run, inspectable afterwards.

The usual complaint about agents is that nobody can tell what they did. That is a tooling problem, and it is solvable.

  • No credit card required
  • Free plan to start

Why agent debugging is hard

  • A final answer says nothing about how it was reached.
  • Failures are hard to reproduce after the fact.
  • Cost is invisible until the invoice.
  • "It worked yesterday" is unfalsifiable without a record.

What every run keeps

The full step history

Each step with its inputs, outputs, and a screenshot, retrievable per run.

Token accounting by stage

Where the spend went — planning, observing, extracting — not one opaque total.

Live while it happens

Stream the run's events and watch the browser in real time.

Failures name themselves

Terminal runs carry an error code and a reason, so retry logic can be deterministic.

How you'd call it

A goal and a schema over HTTP — no selectors to maintain and no browser fleet to run.

“GET /v1/runs/:id/steps and /v1/runs/:id/events to reconstruct exactly what happened, including screenshots.”
The same data the dashboard renders is available to your code.

Questions people ask

How long is run data kept?

Runs, events, and artifacts are retained per your workspace's retention settings and are retrievable through the API until then.

Can I stream events live?

Yes — events are available as a stream, and the live browser view can be watched or embedded.

Is token cost visible per run?

Yes, broken down by stage, along with how much was served from cache.

What will you hand over?

Start with one task. Watch it run, take the wheel whenever you want, and keep the ones that earn their place.