2 October 2026

“Cannot reproduce” to root cause and fix in one working day

How an agent team turned a ticket marked for closure into a root cause, a fix and proof

Press → or space to advance ›
Before: 53 days

One crash, seen once, and no way to make it happen again

Ticket comment, 18 August

I tried to reproduce the crash locally using the same 130-periodicfileupload.t Cram test on a prplOS freedom board setup.

Ticket comment, 30 September

can we [close] it, it hasn’t been seen lately, cannot be reproduced, the corrupted stack is corrupted.

Ticket comment, 2 October

Unless you disagree, I will close this ticket as 'Rejected - Not a problem' in a week.

9 to 10 AugustAn automated test on a Freedom board crashes once. The ticket opens the next day.
8 SeptemberThe saved files of the crashed test run expire.

This is what a careful engineer can do with one crash, no saved files and no way to trigger it. The limit was the process, not the person.

Calendar time, drawn to scale

53 days of waiting against 7 hours 35 minutes

Ticket open, no reproducer10 August to 2 October 2026
0 days
Agent team2 October, 07:55 to 15:30 UTC
0 min

The agent bar, zoomed in 168 times

09:42crash site named
10:20crash on demand
14:09merge requests open
15:26fix shows 0 crashes

Calendar time, not working time. Nobody worked on the ticket every day for 53 days. The agents had a human lead for the whole day.

2 October 2026, all times UTC

One working day, with a human at every gate

07:42
09:00
10:00
11:00
12:00
13:00
14:00
15:00
09:42Crash site named from the one saved crash record
10:20Reproducer: crashes on every run
11:37First fix ready
12:36Reworked fix ready
14:09Three merge requests open
15:26Fixed build: 0 crashes in 1,650 test runs
07:42The ticket gets its closure notice
07:55Planning interview: 36 questions, answered between other work
12:17Returns the first fix: “overgenerated”
14:05Approves the reworked fix
15:27Approves the comment on the ticket
Human decision Agent milestone Start 07:55, close-out 15:30: 7 h 35 min
What broke

The plugin used memory after it gave it back

Think of a coat check. You leave a coat and get a ticket. Someone returns your coat early, and the hook gets a new coat. Then your ticket comes back, and the clerk hands out the wrong coat.

Coatthe record of one file transfer Ticketan upload request that waits for the server Hooka slot in memory

Turning off a transfer gave its record back while an upload request still pointed at it. When the server replied, the plugin read the slot. On the Freedom board, other data already sat there, so the plugin followed a garbage pointer and crashed.

The unsafe window is a few milliseconds. That is why the crash was so hard to repeat.

Memory in use free given back Transfer recordfile name, settings #&?! 0x7f3a…other data points at the record Upload requestwaits for the server reply deleted with the record Upload serveron the network reply nobody waits: dropped
CRASH
A transfer is on. Its upload request points at the transfer record.
Proof

The crash now happens on demand, and the fix stops it

Replay Measured results from QEMU, a virtual prplOS device. Each square is one test run. Totals and the first crash in each row are measured. The order of the other crashes is random.
Current prplOS image
Crashes
Same image with the fix
Crashes
Turn off during an uploadserver holds the reply
0of 0
0of 0
The four steps of the automated testas on the Freedom board
0of 0
0of 0
Random on and offupload every second
0of 0
0of 0
Controlno on and off
0of 0
0of 0
20 of 20runs crash at the first try under a memory checker on a PC.
14 crashesstop at the same instruction as the crash on the Freedom board.
3 new testsone per fix commit. Each fails before its commit and passes after it.
Not field ratesThe test server holds each reply to widen the timing window.
The team

Models from different vendors check each other’s work

gates and decisions tasks delegates change findings backup HumanLead engineerDecides every gate Supervisor · AnthropicClaude FablePlans, rules, keeps the record Executor · AnthropicClaude Opus 5.5Analysis, test harness, fix Helpers · AnthropicClaude Sonnet, HaikuSearch, build plumbing Reviewer · OpenAIGPT-6 astraReviews analysis and fix, ships no code Backup · OpenRouterGLM 5.3, DeepSeekStanding by, barely used
Why a backup?

AI models sometimes refuse tasks that look like attack research. Crashing a program on purpose and studying its memory errors is also the first step in writing an exploit. So a model can decline even when the work is defensive.

In an earlier project, one reviewer model refused to probe a live crash. We keep a backup with models from other vendors, so the work does not stall. The same rules and human gates apply to it. In this project, no model refused.

The reviewer is never the authorA model never reviews its own work. Here, an OpenAI model reviewed the work of Anthropic models.
5 real issues foundReview rounds found five issues in the first fix, including one regression that the executor had added.
A simpler triggerThe reviewer saw that a queued reply is enough to crash. This made the reproducer crash on every run.
Control

A human stayed in charge at every gate

  • Before any code ranA planning interview of 36 questions over about 88 minutes, answered between other work. The answers became the rules of the campaign.
  • 12:17, the first fix goes backTwo AI review rounds had passed it. The human lead read it and returned it as “overly defensive and overgenerated”.
  • 14:05, the reworked fix is approvedOnly then did the agents open the merge requests.
  • 15:27, the ticket comment is approvedThe text went out on the human lead’s word, not before.
  • After the merge request opened, line coverageThe human lead asked for it. It showed 3 changed lines that no review had flagged.

The human catch

Returned

Lines of test code

First fix
about 700 lines
Rework
372 lines
21 minfor the rework
about 8 USDagent cost of the rework

The human lead also found internal campaign labels in the test comments. The rework removed them. The final product code change is 54 lines added and 4 removed, in 3 files.

Cost

Agent time for the whole case: about 171 USD, at most

about171 USD

At API list prices: 133.88 USD recorded for the Anthropic models, at most 33.19 USD for the OpenAI reviewer, and about 4.21 USD for a known gap in the counting.

Human time: the planning interview ran for about 88 minutes. The lead engineer answered between other work, often from a phone, and made each gate decision the same way. Active human time was not tracked.

By phase and model, USD at list prices
Setup
17.48
Analysis
8.39
Reproduce
28.77
Fix
43.96
Review
9.63
Merge requests
7.06
Verify the fix
3.69
Close-out
14.90
Counting gap
about 4.21
OpenAI reviewer
at most 33.19
The reviewer ran on a flat-rate plan, so it cost no extra money.

The deck gives no cost for the human path. Nobody measured it, and we do not guess it.

What happens next

A merge request is not the end

We are here
2 October 2026
1Reporter confirms the fix
2Maintainer review, with line coverage
3Merge requests: component, package list, prplOS
4Automated tests on hardware
5Merge
6Release tag
7Version bumps in prplOS
8Ticket closed
Component merge requestThe fix and its tests. Waits for review.
Open
Feed merge requestThe feed is the package list that prplOS builds from. This request moves it to the fixed version.
Draft
prplOS merge requestPulls the new package list into prplOS builds.
Draft

Devices do not have the fix yet. They get it only when a prplOS build includes the new version.

Review continues

The review after the merge request found a gap

Human decision: the lead engineer asked for line coverage in the merge request, so reviewers can see which lines of the fix the unit tests run. A one-line change from the upstream project turned it on.

The 30 changed lines of the fix, one tick each0 of 30 run by the tests
run by a unit testnot reached by any test

The 3 lines that no test reaches sit on two rare paths: a hook script (a helper program around an upload) that ends or times out, and a reset in the server reply handler.

No earlier review, by AI or by a human, had flagged them. They are now a review item. The work is not done when the merge request opens.

Merge request status, 2 October
  • PassBuilds
  • PassUnit tests
The proposal

Run this as a normal process, steered by Jira labels

A label is a tag on a Jira ticket. Anyone can filter tickets by it.

New ticketenters the queue Agents check firstcan agents handle it? agentic-candidatetag on the ticket, with a reason Maintainer addsagenticthe only way to start the team Agent team, models from more than one vendor Reproduce Root cause Fix and tests Cross-vendorreview Draft mergerequest chain Report onthe ticket Hand back: “I cannot handle this one, and here is why.”Possible at any step. The ticket returns to the normal queue with the reason written on it. Works across vendors: one model’s refusal does not stop a ticket. A person approves the fix and the report.Agents never merge.
Where to start

Start with the crashes that people would close

Then any ticket that an agent can handleagents propose, a maintainer decides
Crashes with no reproducertickets that end as “cannot reproduce”
Start here

The cost of trying is known

At API list prices, this case cost about 171 USD of agent time, at most. A failed try costs the agent spend plus the time a maintainer needs to review it. The ticket then goes back with a written reason.

The gain is a bug that does not stay hidden

A crash closed as “cannot reproduce” stays in the product. Here, the virtual device crashed in up to 76 % of runs once the timing window was widened.

Widen step by step

Add ticket types when maintainers trust the results. Track the cost per ticket, the hand-back rate, and the share of fixes that maintainers accept.

Limits, stated plainly

What this case does not prove

  • No run on a Freedom boardThe proof comes from QEMU, a virtual device, and from a PC. The original crash was on a Freedom board, and no hardware jobs ran.
  • The crash rates are not field ratesThe test server held each reply to widen the timing window. In the automated tests on real hardware, the crash was seen once.
  • Devices do not have the fixThe component merge request is open. The feed and prplOS merge requests are drafts.
  • The fix has one known limitNo freed memory is read, and there is no crash. If a transfer is deleted and a new one reuses its memory slot, a forced upload started within milliseconds can show as finished early. Turning a transfer off and on again is not affected.
  • Active human time was not trackedThe planning interview ran for about 88 minutes. The lead engineer answered between other work, often from a phone, and made each gate decision the same way.
  • The cost is an upper boundThe total of about 171 USD uses API list prices. The OpenAI reviewer ran on a flat-rate plan, so its share of at most 33.19 USD cost no extra money.
  • Unit tests do not yet reach 3 of the 30 changed lines.
Read the evidence

The ticket and the three merge requests

To open these links, you need a prpl login for Jira and GitLab.

The ticket was marked for closure at 07:42 UTC. By 15:27 UTC the same day, it had a root cause, a fix in review and test results. Your turn: which tickets should get the agentic label first?

Build 9ab078e 2026-10-02