CodeMower.comQuality and velocity for AI-assisted development

Code Mower CLI is v1.x; CodeMower.com v1.1 follows an independent release line.

Illustrative Impact and Quality path

A labeled sample of the signed-in CodeMower.com reading path after opt-in metadata uploads. These numbers are illustrative, not a live cohort benchmark.

Illustrative sample data

Code Mower CLI v1.x collects metadata locally. CodeMower.com v1.1 is a CodeMower.com product milestone. This page shows the product shape before sign-in; live dashboards use your team's own uploads.

Reading path in the signed-in product

PRs merged

4

Complete merged PRs in this illustrative window.

Median cycle time

2h

Median open-to-merge time for those PRs.

Blockers caught

2

Reviewer catches that blocked before merge.

Cost per PR

$1.05

Total reported spend $4.20 on 4 complete-cost merged PRs. Missing spend is not treated as zero.

Useful blockers

2

Distinct from advisory comments.

Advisory findings

7

Non-blocking findings kept separate.

False positives

1

Known-clean controls the lane still flagged.

First-pass passes

7 / 9

Useful blockers stay distinct from advisory findings. Escaped-defect rate is not shown.

What uploading metadata unlocks

Reviewer value

Which provider/lens pairs catch useful issues without noisy false positives on your own repos.

Cost and latency

How much each lane costs and how long it takes before you promote it into a stronger workflow.

Data provenance

Dogfood, imported history, and calibration evidence stay visibly separated so the dashboard does not overclaim.

Enable-next guidance

Practical recommendations for informational, selective, or merge-gating lanes as evidence accumulates.

Evidence classes stay separate

CodeMower.com keeps operational dogfood, imported history, and reviewer evidence visibly distinct so the dashboard can be useful without overstating what the data proves.

Current dogfood

4

steps

54

events

Use for ingestion health, recency, repo coverage, and whether dogfood workflows are running.

current repo metadata and provider inventory

Imported history

2

steps

18

events

Coverage and timeline context before dogfood was enabled.

sanitized GitHub Actions history

Reviewer evidence

3

steps

66

events

Use for reviewer and lens value once verdicts are tied to known-clean and known-blocked cases.

metadata-only reviewer verdict artifacts

From Impact to Lane Intelligence

Impact and Quality answer whether the loop worked. Lane Intelligence is what CodeMower.com calls the next step: turning that same metadata into a builder x reviewer pairing per work type, scoped to one repository. The illustrative pairings below use a ten-run sample that is large enough to route and a two/three-run sample that is deliberately too small, so the "collect more evidence" outcome is visible, not hidden. Confidence and sample size are never blended into one score, and there is no cross-customer benchmark; enabling a lane stays an operator decision made outside this dashboard.

Builder x reviewer pairings

Work-type recommendation matrix

Conservative builder/reviewer pairings by work type from direct evidence. Builder and reviewer metrics stay separate; actions are descriptive only and never change repository policy. Metadata-only evidence; no source, diffs, or secrets.

Open reviewer value

Recommendation rail

One line per work type: suggested pairing, action, confidence, and sample sizes.

  1. backend-risk · codemower-ai/code-mowercursor / cursor (gpt-5) + codex / codex-audit (gpt-5)
    high · b:10 r:10Route
  2. docs · codemower-ai/code-mowerno eligible builder + no eligible reviewer
    insufficient · b:2 r:3Collect more evidence

Pairing matrix

Same decision detail as desktop: pairing, action, confidence, samples, limitations, alternatives, and backing evidence.

backend-risk

codemower-ai/code-mower

Route

high confidence

cursor / cursor (gpt-5)

first-pass 83% · 10 PRs

exec cost $1.10 · cycle 5400s

high confidence

codex / codex-audit (gpt-5)

useful 80% · 10 runs

cost $0.4200 · selective-trigger-candidate

high confidence

Route backend-risk in codemower-ai/code-mower to this pairing. Builder evidence [cursor / cursor (gpt-5): first-pass 83% over 10 PRs, execution cost $1.10]. Reviewer evidence [codex / codex-audit (gpt-5): useful 80% over 10 runs, cost $0.4200 (selective-trigger-candidate)]. Descriptive only; enabling or gating lanes stays a manual repository decision.

Evidence limitations

None recorded.

docs

codemower-ai/code-mower

Collect more evidence

insufficient confidence

No eligible builderNo eligible reviewer lane

Collect more evidence before pairing docs in codemower-ai/code-mower. Builder evidence [no eligible builder]. Reviewer evidence [no eligible reviewer lane]. No winner declared.

Evidence limitations

  • builder sample too small (2 PRs, needs 5)
  • reviewer sample too small (3 runs, needs 5)

Competing alternatives

  • cursor / cursor (gpt-5) (builder, n=2, low; first-pass 50% over 2 PRs)
  • claude / claude-audit (claude-opus-4) (reviewer, n=3, insufficient; useful 33% over 3 runs (collect-more-evidence))

Issue planning lineage

GitHub issue to delivery

codemower-ai/code-mower#269

Code Mower can carry metadata from a GitHub issue into a posted plan, work order, builder/reviewer run, PR, merge, and upload. The dashboard shows what is captured and what is still missing before claiming end-to-end delivery evidence.

4/7 steps captured

Next missing evidence: Builder.

Signed-in dashboard shape

Supporting illustrative chrome from the same sample session. It is not additional live evidence and does not change the Impact and Quality numbers above.

Cloud activity

Benchmark signal at a glance

Recent uploads, reviewer events, and source provenance for your team. The goal is to make dogfooding health visible before anyone has to read the raw tables.

Last 30 days

Uploads

41

Metadata bundles received for your teams.

Evidence events

110

18 provider inventory snapshots tracked separately.

Repos

4

Rows in repo rollups for the selected dashboard window.

Spend

$42.18

Recent reported reviewer spend.

Latency

84s

Average recent provider/runtime latency.

Activity trend

Upload and event volume over the last 14 days.

UploadsEvents

Latest day

06-15

Latest uploads

4

Latest events

6

Source mix

Dogfood, calibration, and imported history stay separated so the signal is explainable.

code-mower-local54 · dogfood
lens-calibration-corpus40 · calibration
historical-backfill18 · history
github-actions16 · dogfood
Other sources18 · other

Dogfood

54

Calibration

74

Historical

18

Anonymous uploads

0

Reviewer signal

Which lanes are earning trust?

These cards summarize useful signal, noisy lanes, spend, latency, and the next lane recommendation from your own uploaded metadata.

Best reviewer so far

Not enough data yet

Upload calibration events to identify which reviewer/lens has useful signal on your codebase.

Noisy lanes

No noisy lane yet

No recent lane has more false positives than useful signal.

Cost / latency

$42.18 / 84s

Recent reported provider spend and average runtime latency across uploaded events.

What to enable next

Promote codex-audit on risky PRs

Base Codex reviews have the best useful-rate in this sample with low false positives and acceptable latency.

Evidence maturity

How strong is this data?

Code Mower should not overclaim. This ladder shows which decisions are supported now, which data is only context, and what is still needed before treating provider/lens comparisons as benchmark-grade.

Productivity model

Parallel AI review capacity

Code Mower can exceed 24 hours/day because this is aggregate portfolio capacity across agents, reviewer lanes, and repos. It is not a claim that one human can work more than one wall-clock day.

Reviewer runs

66

Provider/lens runs represented in recent metadata.

Useful signals

41

Useful issues or decisions captured by reviewer events.

Capacity modeled

2d 5h

Conservative aggregate review and follow-up attention estimate.

Net value

$7,870

Modeled value minus reported provider spend in the recent window.

Assumptions

Reviewer run20 minutes of human review attention
Useful signal45 minutes of triage/fix follow-up
Loaded hourly rate$150/hour
Reported provider spend$42.18

Calculation

Review capacity66 runs x 20 min + 41 useful x 45 min = 2d 5h
Aggregate pace1.8h/day across parallel lanes
Gross modeled value$7,913
Annualized run-rate$94,444

This is a model, not accounting. v1 should let teams tune assumptions and should distinguish reviewer value, builder value, avoided defects, and cycle-time acceleration.

Calibration cases

24

Known-clean and known-blocked PRs used to judge reviewer behavior.

Reviewer runs

128

Structured reviewer and lens runs across the calibration corpus.

30-day spend

$42

Reported API or subscription spend from local metadata.

Avg latency

84s

Average reviewer runtime for recent structured audits.

Best reviewer so far

codex-audit

Strong useful-rate with a low false-positive rate on the sample corpus.

Noisy lane

gemini-cli / operability

High false-positive rate means this lane should stay advisory until calibrated further.

What to enable next

Selective triggers

Run Codex on every candidate PR; trigger Claude quality lens on backend/auth/data changes.

What this dashboard should help you decide

Who should review every risky PR?

Use codex-audit as the baseline lane.

It caught the known-blocked control and stayed quiet on the known-clean control.

Who should be selective?

Use Claude for matching backend/auth/data classes.

It added useful independent signal, but with higher cost and latency in this sample.

Who should stay informational?

Keep the experimental lens out of branch protection.

It was cheap, but it blocked a clean control and missed a blocked control.

Reviewer value table

LaneLensUseful rateFalse positivesCost / runLatencyRecommendation
codex-auditbase82%9%$0.4271sMerge-gating eligible after one more clean cycle
claude-auditcontext-driven-quality74%14%$0.6196sSelective trigger for higher-risk PRs
gemini-clioperability38%31%$0.0852sInformational until more calibration data exists

Why contribute metadata?

The local OSS tool already gives you private reports. Opt-in cloud sharing adds team history, dashboard rows, token management, export/delete controls, and a place to compare reviewer value over time without keeping every terminal artifact in someone's laptop. Cross-team comparisons come later, after enough teams opt in for the data to be honest.

Ready to try the local path? Follow the setup guide. Want the OSS source first? Open GitHub.