Findings report · 30 July 2026
We harvested 380 ideas from Xero's public Product Ideas board, picked eleven, analysed each one properly, built working mockups, and filmed a concept walkthrough of every one. This is what we found — about the requests, about the board, and about our own accuracy.
Vote count tells you how many people hit a problem. It tells you nothing about what fixing it costs.
This is the finding that survived contact with every one of the eleven. The board presents vote counts as its ranking mechanism, which invites the reading that a high number means "build this next". Sorted by votes and coloured by what we actually found, that reading falls apart.
Read the colours down the length of the chart. Constraints appear at 694 votes and at 142. Clean builds appear at 444 and at 103. The most expensive change in the whole set — a second approval level on bills, which disturbs the status model, the audit story, and every already-approved bill in the organisation — has the smallest vote count here.
That inversion is not noise. It is what you would expect: the number of people who notice a friction depends on how often they hit it, and the cost of fixing it depends on how much of the system it touches. Those are unrelated quantities. A board that ranks by the first and stays silent on the second will keep producing this surprise.
Two buildable as asked, four better in a different shape, five carrying a cost worth stating.
We committed in advance to three possible conclusions and to not softening any of them. The distribution matters less than what each category turned out to mean.
Payer names on a reconciliation row and a default for the Approve button are both small, both well precedented, and neither disturbs anything. The temptation with work like this is to manufacture a complication so the analysis looks rigorous. That would have been dishonest. The Approve default is a settings row and a button label; it happens to sit against an irreversible action that currently costs one click while the safe one costs two.
In every reshaped case the requester had correctly identified the friction and reached for a display change where a behaviour change was needed. Payment method on bills is the clearest: a column is passive, and it tells you something at exactly the moment you are least likely to read it. The failure described is a slip, and slips are not fixed by adding information at the site of the slip.
None of these are unbuildable. Each one has a consequence that the request does not mention and that only appears when you build it. Unapproving a sales invoice desynchronises Xero from a document the customer already holds. Custom fields land in the editor but not on the PDF the customer receives, which is the inverse of the stated purpose. Splitting a batch payment splits its identity, so a remittance advice already sent matches neither half.
Two of the eleven collide with Xero's own committed work. Switching autosave off contradicts moving the Save button to the bottom of the invoice, which is In development with 247 votes — an explicit Save button in an autosaving editor is decorative. And unapprove has already been Accepted for bills, which is what makes the invoice case interesting rather than simple.
The ranking mechanism hides the thing most worth seeing.
One subject was not filed by anyone. We synthesised it from the harvest because the pattern was too consistent to ignore: the board ranks ideas by votes inside each forum separately, so the same request filed in two forums is ranked twice, counted twice, and can be decided two different ways without either decision acknowledging the other.
The obvious fix — merge duplicates and sum the votes — is wrong, and it took the longest analysis in the set to establish why. A merged total claims an endorsement nobody gave. Nothing decides "same idea" reliably except a person: a lexical scan over all 380 titles returns roughly fifty candidate pairs of which about six are genuine. Worse, that scan's own scores rest on invisible judgement calls. Treat one unremarkable word as a stopword and the clearest true duplicate on the board drops below the threshold and is missed.
What has value was never the arithmetic. It is seeing that the same request was filed twice and answered two different ways — which the current board makes structurally impossible to notice.
We audited every factual claim before publishing. The data was clean; the arithmetic on top of it was not.
Roughly 190 atomic citations were checked — every vote count, status, title, URL and forum attribution, plus 36 verbatim quotations. Zero were wrong.
Of 38 derived claims — sums, ranks, family sizes, percentages — ten were wrong. Six of eight structural claims about the repository were wrong. Four superlatives were wrong and six more were unsupportable at the stated scope. The single most damaging error appeared three separate times: one idea was described as the highest-voted in the harvest when it is second.
The lesson generalises past this project. Copying a number accurately is easy and we did it 190 times without a slip. Adding numbers up, ranking them, or calling one of them the largest is where the errors live — because each of those silently introduces an assumption about scope that nothing checks. Every family total we report is a lower bound, and no board-wide superlative is supportable from a 380-idea sample of roughly 6,000.
One unexamined assumption nearly cost every film 22 seconds.
The films are made with a scripted browser-filming engine. It offers five beats that put words on screen. All eleven films were written using one of them.
Because nothing enumerated the available surfaces, setup narration went into the caption bar — the single worst choice available, because captions are the one surface with a hard word limit and the one whose duration scales with text length. The project was one decision away from either lengthening every caption past its limit or adding 60% more of them.
| Approach | Mean film | Longest | Craft floor |
|---|---|---|---|
| Setup in captions | 87s | 113s | Repeated warnings |
| Setup on a title frame | 65s | 75s | Clean |
| As shipped | 67s | 79s | Clean, strict mode |
Understanding one beat we already had was worth 22 seconds per film and saved the caption register entirely. The failure was not a missing feature. It was not knowing what was already there.
Each links to its own page: the friction in the requester's words, what the walkthrough shows, and the reasoning in full.
| Request | Votes | Board status | Assessment | Film |
|---|---|---|---|---|
| Payer names instead of “multiple items” | 103 | submitted | As asked | 61.4s |
| Select default for the Approve button | 444 | Not in pipeline | As asked | 58.3s |
| Payment method on bills | 222 | submitted | Reshaped | 69.3s |
| Add subtotals | 348 | Not in pipeline | Reshaped | 63.6s |
| Search across all bank accounts | 250 | Not in pipeline | Reshaped | 61.7s |
| Demand split across forums | — | Not on the board | Reshaped | 68.2s |
| Multiple levels of approval for bills | 142 | submitted | Constraint | 78.7s |
| Split a batch payment when reconciling | 345 | Not in pipeline | Constraint | 70.6s |
| Option to switch autosave off | 480 | Not in pipeline | Constraint | 68.9s |
| Unapprove option | 518 | Not in pipeline | Constraint | 71.6s |
| Custom fields on invoices and contacts | 694 | Not in pipeline | Constraint | 68.1s |
The number is a floor, not a mandate.
Every film states this before it uses the number for anything, and it is worth repeating here. A vote count is a lower bound on how many people hit a friction. It is not a measure of how much the friction matters, whether the requested form is the right one, or whether the cost of building it is worth paying. Those three judgements belong to whoever owns the product, and the whole point of filming these is to put the request and the constraint in front of that person at the same time.