Skip to content
trawler
Use Trawler

Core concepts

A few words carry most of the meaning in this product.

Where everything you create lives: projects, runs and the one model key that pays for them. Your first sign-in creates one for you and makes you its owner, unless your email address has a pending invitation to another. Workspaces are kept apart by the database itself, so one workspace cannot read another’s projects, runs or key. See Sign-in and your workspace.

A project is one product under test: the address it starts from, its name and description, and its plan — the people, their goals and any test accounts. Setup proposes the plan from the product’s page; the plan page is where you change it. See From a URL to a plan.

A simulated user: a first name and a few sentences saying what they are like, written to them in the second person — their situation, their patience, what they care about. The briefs setup writes describe the person and never the product; where things are is for them to find out. A person can be given a test account to sign in with.

Something a person wants to get done in the product, phrased as the outcome rather than the steps — invite a teammate to the workspace, not click Settings, then Members. Every person pursues every goal of the plan, in order, and marks each one reached or failed.

One execution of the plan, with a model, a cap and a number within the workspace (Run 0001). A run takes a copy of the plan when it starts, so editing the plan afterwards does not change a run that has begun — except that a test account removed in the meantime is gone for the sessions and replays still to come. A run moves through four stages: Use, Replay, Judge and Report.

What a person reports while using the product, recorded the moment they see it:

Kind What it claims Replayed
Defect Something about the product behaved wrongly. It needs at least two steps, and people are told to write them literally enough for a stranger to follow — exact addresses, labels and values — with what went wrong kept out of them. Yes
Friction Something about the person’s experience: they could not find something, or it was not clear. Its steps are the path they actually took. No

Each finding carries a severity — low, medium or high — and the goal the person was working on.

After every person has finished, each defect is replayed: a fresh agent in a new browser gets the product’s address and the steps — and the test account of the person who reported it, if they had one — but not the title, not what the person saw, not the goal. It reports whether it could carry out every step and what the page showed. A judge then compares that report with the original claim and answers confirmed, refuted or inconclusive; a judge that gives no answer leaves the defect could not be judged, which you can judge again once the run is over, if its cap has room left. See Reading the report.

Where a run stops spending on model calls, in dollars — $2 unless you choose otherwise, and never more than $50. Spending is counted for every call the run makes; once it reaches the cap the run stops and keeps what it found, though the last call can take it slightly past. A run on a model with no known price stops after 3 million tokens, and only through OpenRouter, which reports each call’s cost, is it also held to the cap. See Models, estimates and the cap.