Skip to content

Factories > Measure and improve

Measure and improve a factory

Open in ChatGPT ↗
Ask ChatGPT about this page
Open in Claude ↗
Ask Claude about this page
Copied!

An overview of how Warp Factories measures and improves a factory, with links to Scorers, Self-improvement, and Benchmarks.

Warp Factories tracks what your factory produces and how well it performs, so you can spot a problem, test a fix, and decide whether to keep it.

FeatureWhat it tells you
Dashboard metricsHow much work the factory produced, and what it cost.
ScorersWhether completed runs meet criteria you define.
BenchmarksHow different configurations perform on the same tasks.
Self-improvementWhich repeated failures get investigated and turned into follow-up work.

The Dashboard page shows activity, cost, autonomy, and evaluation results:

MetricWhat it shows
Total runsAll agent runs, with breakdowns by agent type, status, source, model, and more.
PRs openedPull requests created from factory work, counted once, in the period they were first observed.
PRs mergedOf the PRs opened in a period, how many later merged. Opened and merged draw from different data sources, so a period’s merged count can occasionally read higher than its opened count for that period.
AutonomyThe share of the factory’s merged PRs that needed no human code push before merging. Opening the PR counts as a push, so a human-authored PR that a factory run later revised doesn’t count as autonomous.
PR cycle timeThe median time the factory’s merged PRs took from run kickoff through PR, first review, and merge, with a median for each stage.
Cost per PRThe median cost of PRs opened in the range, split evenly across a run’s PRs when one run produces more than one. View it broken down by cost component or by PR size (S/M/L/XL, split at 100/500/1,000 changed lines).
Most expensive PRsThe highest-cost pull requests.
Scorer cardsResults from your Scorers.
Self-improvement PRsThe three newest Self-improvement pull requests, regardless of the selected date range.

Cost per PR is an estimate, not a billing figure: it counts recorded credits and can undercount actual usage.

Use the Dashboard page to pick which runs to investigate, not to conclude what caused a change. Total runs includes evaluation, benchmark, and Self-improvement runs, so a higher run count with a flat PR count could mean harder tasks, retries, or measurement activity.

A Scorer uses an LLM judge to classify completed runs against criteria you write — for example, “did the agent run the tests before opening a PR?” See Configuring Scorers for its fields and how automatic and on-demand scoring work.

A benchmark compares model and runner configurations for a single agent on the same fixed tasks. Use it to test a configuration change before you apply it to production. See benchmarking factory agent configurations for the workflow.

Self-improvement turns a Scorer’s repeated failures into follow-up pull requests, against application code or the factory’s own definition. See Configuring and reviewing Self-improvement for how to turn it on and review its pull requests.

Change one measurable thing at a time:

flowchart LR
  Define[Define a Scorer] --> Baseline[Collect a baseline]
  Baseline --> Inspect[Inspect failures]
  Inspect --> Benchmark[Benchmark a candidate]
  Benchmark --> Adopt[Review and adopt]
  Adopt --> Monitor[Keep monitoring]
  Monitor --> Inspect
  Inspect -.->|Repeated failures| Improve[Self-improvement]
  Improve -.-> Adopt
  1. Define a Scorer. Pick one agent and one failure mode you can observe. Write the judge instructions and classifications, then score a few runs manually and compare the judge’s results against your own review.
  2. Collect a baseline. Let automatic scoring run until results reflect normal work. Record the Scorer settings, date range, and relevant costs.
  3. Inspect failures. Read the judge’s reasoning and the underlying runs. Look for causes like missing context, unclear instructions, or missing tools. Turn on Self-improvement when the same failure keeps repeating.
  4. Benchmark a candidate. Compare configurations of that agent on the same tasks, with enough repetitions to trust the difference.
  5. Review and adopt. If the evidence supports the change, make it. Review Self-improvement pull requests with the same standards as human-authored ones.
  6. Keep monitoring. Leave the Scorer active and compare new results against your baseline. Revise the Scorer, or set its sample rate to 0, when its criteria no longer match what your team needs.