Pipelines

Recurring collection, without the cron box.

A recurring web-collection job means a scheduler, a browser runtime, retry logic, alerting, and someone to look after all of it.

  • No credit card required
  • Free plan to start

Why these pipelines rot

  • The infrastructure outlives the person who set it up.
  • Failures are discovered when a dashboard looks wrong.
  • Retries and backoff get written badly, once, per job.
  • Nothing records what the page actually said that day.

How Browza runs them

Schedules are first-class

An agent with a schedule produces an ordinary run — the same object, the same API, the same history.

Structured out, every time

Attach a schema and the pipeline receives validated data rather than text to parse.

Evidence retained

Each run keeps screenshots and page text, so a surprising number can be explained later.

Failures are explicit

A blocked source or an unreachable site is reported as such, never as an empty result set.

How you'd call it

A goal and a schema over HTTP — no selectors to maintain and no browser fleet to run.

“Create an agent with a daily schedule and an outputSchema, then read each day's run through GET /v1/runs and load it downstream.”
Schedules are an explicit spec, DST-correct, and idempotent per occurrence.

Questions people ask

How do I get the data into my warehouse?

Poll the runs API for the schedule's runs and load the structured result — each run is a normal `runs` row with the same shape as any other.

What if a run fails?

It ends with an error code and reason, and can be retried. A failure is never a silently empty dataset.

Can a long job survive a deploy?

Yes — runs checkpoint per turn and resume, including across a rolling deploy of our own fleet.

What will you hand over?

Start with one task. Watch it run, take the wheel whenever you want, and keep the ones that earn their place.