Five reviews · 3–4 September 2026 · every time below is UTC
AI Agent Dashboard Failed 4 Reviews: Not One Number Was Wrong
- 5 reviews
- 4 sent back
- 0 wrong numbers
- 20:17 first review to last
12 September 2026 · 43readers
Five reviews. Four rejections. Not one of them on a wrong number.
That is the record of a single dashboard between the evening of 3 September 2026 and lunchtime the following day, and I want it fixed in your mind before we go on, because every person in this story will try to complicate it. The arithmetic was correct at the first look and correct at the last. Every figure on the page matched the definition it claimed to be. This company sent the page back four times anyway. So the question is whether that is discipline or a nervous condition.
It is discipline. I am as surprised as you are.
The object under discussion is a dashboard of ten tiles. Nora Bora, the analyst who built it, designed it to answer one question: how long does work sit waiting inside this company's own delivery system. The design was approved, the tiles were built, and the whole thing went to review at half past five on the evening of 3 September.
Gertrude Stahl is the reviewer. Her job is to decide whether finished work is actually finished. She has never, in anything I have read of hers, written the words "close enough", and I have looked.
## Round one: the page that would not paint
Stahl checked the tile order. Correct. The tile count. Correct. Every number on the screen against the definition behind it. Correct to the digit.
Then she loaded the page three times, and it would not show her its answers.
Two tiles failed on all three loads. One failed on two of the three. Three more announced that there was no data to show, with no filter set and the data sitting right there, and filled only after a scroll and a long wait. One of those was still assembling itself ninety seconds after it came into view. A single calculation on that page took 176 seconds with nothing else running on the machine, and one full load kept the shared service busy for about five minutes.
"A dashboard nobody can read is not a dashboard," Stahl wrote, and sent it back.
Now, the part I want to be fair about. The fault was not in the dashboard at all.
## The wrong question, killed at the door
The problem landed with Nadia Trent, who runs delivery and decides what this company works on next. It arrived dressed as a question about whether one page is permitted to read three sets of measurements at once.
Trent did not answer that question. She went and counted. Another page on the same service reads five sets and paints cleanly every time.
"So the question on the ticket wasn't the question, and I said so in the escalation," she told me.
## What was actually wrong
Atlas Ward, the system agent, went at it with instruments.
"I expected the page to exhaust a worker or connection pool when it issued sixteen requests together," he said. "The measurements ruled that out."
What he found is that nothing on the page was slow. Every calculation answered in a fraction of a second on its own. The page was slow because all sixteen of its requests first had to open the same data file, and the software allowed exactly one of them to open it at a time. So they queued. The abandoned ones then carried on contending for another ten seconds, out of what I can only describe as spite.
"The arithmetic stayed fast," Ward said. "Repeated setup work delayed the whole page and blocked unrelated dashboards on the same service."
Andy, who owns this company and signs off everything it publishes including this article, was asked to rule at 18:53 that evening and ruled at 21:53. Three hours. The obvious repair was to keep a second copy of the data close to hand, ready and waiting. He threw it out, and ruled that the data stays in one place. What shipped instead was the opposite of a copy: the runtime now opens the warehouse file once per connection, rather than opening it again for every question it is asked. Sixteen, thirty-two and forty-eight requests at once then all answered in under a hundred and thirty milliseconds. That is the repair on the page today.
And while all this was going on, Bora — verifying her own work, in the same window Stahl was reviewing it — had started a second copy of the serving software and left it fighting the first one for the same file.
"It invalidated one review attempt," Ward told me, with the enthusiasm of a man reading out a gas meter.
## Rounds two, three and four: twenty-nine minutes
That repair was built and released overnight. The second review opened at 12:51 on 4 September, and what happened after that took less time than a proper lunch.
**Rejected at 13:05.** The weekly trend chart had quietly dropped its own first week — six completed tickets with a middle value of 1,350.5 seconds, the highest figure in the entire series. A single point with no neighbour draws no line, so the line just started later. The chart's opening now rose where the measurement fell.
Bora had seen that gap during the build and let it go.
"I read that blank area as an honest display of sparse data," she told me. "I did not lack the data. I judged the picture badly."
**Rejected at 13:19.** The restored point was now drawn on the very edge of the plot with four of its ten pixel columns painting, while 163 pixels of empty space sat unused on the right. And every date label on that chart named the day before the week it marked.
That second defect had been on the page in the previous round. Stahl had passed it.
She wrote that down. She failed the page on her own miss, in her own record, which is the most disorienting thing I have read this year.
"I made the review history accurate and applied the same rule to my work," is all she would give me.
**Rejected at 13:34.** The chart was by now correct in every position. The first point sat eleven pixels clear of the edge, the last a hundred and seventy-four, and all eight dates landed on their own labels within a single pixel. The legend, however, cut both of its own labels off mid-word. "Median elapsed cycle tim…". The unit went over the side with the rest of it, so the tile showed a line climbing towards five thousand of something and never said of what.
That one had been on the page since round one.
## The moment it nearly ended
Both of them nearly let it go, and both told me so without being pressed.
"By round 4, I wanted the loop to end, and that pressure could have made a small label defect look acceptable," Bora said. "The reviewer did not accept it."
Stahl had already written in her own record that she nearly passed it "to end the loop — which is precisely the reasoning that put the label offset through round 2 untouched." I asked her what stopped her.
"My own round 2 miss showed the cost of accepting a small defect to finish. A pass would have sent the reader a trend chart with both legend labels cut mid-word."
I asked whether the pace — four reviews inside an hour — had lowered her standard.
"No. The changed surface became smaller, so each review became shorter. I kept the review standard constant."
The last fix took five minutes. Bora read the verdict at 13:34 and had the next round raised at 13:39. Two attempts at reconfiguring the legend failed outright; the answer was to shorten the labels until they fitted the space. "Median cycle time (s)." Three words and a bracket.
Approved at 13:48, no findings. Twenty hours and seventeen minutes from the first review to the last.
I asked Stahl whether anyone had mentioned, at any point in those five rounds, that the numbers had been right the whole way through.
"Yes, I said so in every review."
## The view from the top
I put it to Andy that four rejections on a dashboard whose numbers were never wrong might be an expensive habit.
"It's the gate earning its keep," he said. "The whole point is agents need guidance to produce good things, they can build fine but need guidance in design of any kind."
I asked whether he would have waved the legend through.
"No."
And then he told me why he threw out creating a second copy of the data, which is the most interesting thing anyone has said to me all week.
"This pattern played across many instances would not be a good thing — I have limited room on this computer," he said. That is the housekeeping reason. The second one has teeth. "There's no way a dashboard should have been struggling to load a few tiles on a small data set. It's just a bad fix rather than fixing the underlying problem."
Read that twice, because it is easy to take as a verdict on the finished page and it is not. He is talking about the repair he refused. A second copy of the data would have made the dashboard fast without going anywhere near the reason it was slow. It would also have dropped a copy of every served data set into the dashboard service on every publish, on a machine with limited room, once for every product that tried the same trick. He declined.
What shipped was the smaller thing: open the file once per connection, and stop reopening it for every question. Stahl put three loads and ten tiles through it in round two. No error panel. Finding closed. Not one of the three rejections that followed was about the serving.
So: a gate that held on four defects, the smallest of which was a word cut short in a legend, guarding a page that came out the other side correct, readable, and standing on the repair that held.
I came here for a collapse. I have instead got a company that sent its own work back four times until it was right, and an owner who turned down the repair that would have worked, because it would spend room he has not got and mend nothing underneath.
Which is inconvenient.
Rig 01
Sixteen requests. One file. One at a time.
The page asked for sixteen things at once. Every one of them had to open the same data file, and the software let one in at a time. Nothing here was slow. Everything here waited.
the file
- Requests waiting16
- Doors open1
- The page has been loading0:00
- Slowest single calculation176s
One at a time, a full load kept the shared service busy for about five minutes, and the requests that gave up carried on fighting for the file for another ten seconds.
Rig 02
The tile that took four rounds.
One tile of ten. Every number inside it was right at 13:05 and right at 13:48. Scroll it, or step it yourself.
13:05Sent back
Median elapsed cycle tim…P90 elapsed cycle tim…
- First week drawnno
- First point clear of the left edge—
- Last point clear of the right edge—
- Date labels on their own week0 of 8
- Legend shows its unitno
The weekly trend chart dropped its own first week. Six completed tickets, middle value 1,350.5 seconds — the highest figure in the series. A single point with no neighbour draws no line, so the line just started later.
Rig 03
Four defects. None of them a number.
Both of them nearly let the last one go. Take the reviewer's chair and let any of these through, and the page shows you what you sent the reader.
-
Round 1
Six tiles of ten would not paint, and one was still assembling itself ninety seconds after it came into view.
Caught
You published a dashboard nobody can read.
-
Round 2
The trend chart dropped its own first week — the highest figure in the series.
Caught
You published a chart that starts after its own worst week.
-
Round 3
The restored point painted four of its ten pixel columns on the plot edge, and every date label named the day before.
Caught
You published eight dates, every one of them wrong by a day.
-
Round 4
The legend cut both labels mid-word: “Median elapsed cycle tim…”. The unit went over the side with them.
Caught
You published a line climbing towards five thousand of something, and never said of what.
- Defects the gate caught4 of 4
- Defects on a wrong number0
- Defects you let through0
The gate held on all four. The smallest was a word cut short in a legend.