telemetry ·ai ·agents ·debugging ·claude-code ·observability

Stop Describing the Bug

9/22/2026

8 minutes read

"It hangs on the iPhone. Almost every session."

That was my bug report, more or less word for word. It was accurate. A few days later the data would confirm it to the turn. It was also useless, and not because the agent reading it was careless. There was simply nothing in it to work with.

The project is Turnwise, a Rubik's cube solver that runs in the browser. You show the camera one face of the cube, turn it the way the on-screen guide says, show the next face, and so on until all six are in. On my Mac's webcam, the scan moved along. On my iPhone, it would capture a face and then sit there, cube held perfectly still in front of it, waiting for something. Pointing the camera at the desk for a second and back again unstuck it. I had no idea why that worked, and neither did the agent.

The agent could read every line of the scanner. It could not hold the cube. It could not see my lamp, or the way my hand drifts during a turn, or feel the three-second pause that turns into fifteen. Everything it knew about the bug came through me, and I was a very lossy channel.

The Tempting Answer

The obvious fix for "the agent cannot see" is to let it look. Agents can take screenshots now. They can drive a browser, click through an app, and describe what is on the screen. It works, in the sense that a straw works for emptying a pool.

A screenshot is one frame with no history. Each look is a round trip through a model: capture, encode, reason, decide what to look at next. A good session gets you a picture every few seconds. And the picture shows what a human would see, which in this case was a cube, a camera preview, and a guide arrow that refused to go away. None of that explains anything. The bug did not live on the screen. It lived in a decision the scanner made about ten times a second, and in the frames between the ones anybody would think to capture.

For scale, here is what the telemetry eventually collected from twenty scan sessions.

Figure 01·What thirteen minutes of scanning looks like once it is data

No screenshot loop comes near that, and even one that did would not know what the scanner decided about each of those frames. The pixels were never the whole story. The verdict was.

Translation, Not Observation

So the job was not to give the agent eyes. The job was to take the experience and translate it into the one medium an agent reads fluently: rows.

That is what scan telemetry became. Every scan attempt on a dev build is recorded as one session. Every analyzed frame is a row: what each of the nine cells measured, how the exposure looked, and what the turn detector concluded. The session is held in memory and sent to a small SQLite database on the dev server once, when the scan ends. The scanner itself does not know it is being recorded; it just publishes what it does, and the recorder listens.

Early on, the first version streamed data every two seconds. I pushed back. Phones hide tabs, drop connections, and kill background requests, and I was not interested in if those edge cases would happen, only when. One complete package at the end of a session removed a whole category of problems instead of defending against each of them.

Figure 02·From a scan in my hand to a question the agent can answer

The first real sessions from the phone arrived complete, about nine and a half analyzed frames per second, with no gaps in the record numbering. For the first time, "it hangs on the iPhone" was something the agent could query.

It still could not answer the question.

Telemetry Grows With the Question

The first version recorded numbers only: cell colors, exposure, the detector's internal state. That turned out to be a subtle trap. Every row described the scan from the point of view of the detector under suspicion. When the agent looked at a stuck turn, it could see the detector waiting. It could not see why, because the only witness in the room was the defendant.

What was missing was ground truth: what was actually in front of the camera. So the second version started recording the camera's own pictures, small enough to be cheap and large enough to be honest. A 32 × 32 crop of the square the scanner reads the cube from, and a 48-pixel thumbnail of the whole frame. Then I scanned the cube again, about twenty times, on three cameras.

This is one of those turns, from the phone, frame by frame.

Sixteen small camera crops of a Rubik's cube face over fifteen seconds. The cube turns from a green face to a yellow-centered face within four seconds, then sits still for eight seconds with every frame reading nine of nine cells. Three frames of the desk read eight of nine, and the face is captured at fifteen seconds.
Image 01·What the scanner saw during one stuck turn, with how many cells it read in each frame
The same sixteen moments as whole-frame thumbnails: a hand holding the cube steady in front of a keyboard, then the camera pointed at the desk, then back at the cube.
Image 02·The same moments, whole frame: the cube held still, then the camera pointed at the desk

Look at the frame at three seconds. The cube is visibly mid-turn, two faces smeared across the crop, and the scanner still reads all nine cells cleanly. By four seconds the new face is sitting there, still, perfectly readable. It stays that way for eight more seconds. It is never captured.

The old detector had a simple rule: after capturing a face, wait for three frames in a row that do not read all nine cells, then treat the next steady face as new. On a webcam, a turning cube blurs, and the partial reads come for free. On an iPhone, with its sharp sensor and aggressive processing, a cube in the middle of a turn still reads as a complete face. The partial reads never came. The detector was waiting for a signal the camera would not produce.

And the desk? It read as nine clean cells too, at first. Only a few frames of it came out imperfect, three in a row, and that was enough to re-arm the detector. My "point it away and back" trick was not a trick. It was me, by hand, manufacturing the signal the phone was too good to make.

all 9 cells read (137 frames)fewer than 9 (11 frames)old faceturningnew face, held stillpointed awaybackre-armedcaptured0 s2 s4 s6 s8 s10 sseconds since the previous face was captured
Figure 03·Every analyzed frame of the same turn: 137 of 148 read all nine cells

The pictures did one more thing: they made the fix obvious. The problem was not a threshold to tune. Waiting for three imperfect frames is the wrong question to ask a camera that rarely produces one. The new detector recognizes a face by its colors instead, and captures when the view holds still on a face other than the one it just captured. Before it shipped, it was replayed against every turn in the recorded sessions, offline, with nobody holding a cube.

None of that was in the first design. It could not have been. You do not know what to record until the data you have fails to answer the question you now have. Each failure told us what to add next.

Figure 04·Each extension to the telemetry came from a question the previous version could not answer

The Data Does Not Take Sides

The most useful property of the telemetry turned out not to be what it showed, but whom it contradicted. It did not play favorites.

It contradicted me. I finished a full scan on the phone, told the agent it was in the database, and it was not. Nothing had arrived, and nothing on either end had said so. My memory of sending it was worth exactly as much as the empty table. That gap is why every delivery is now logged on both the phone and the server.

It contradicted the agent doing the implementation. Two end-to-end tests started timing out, and the explanation offered was machine load. The agent reviewing the work did not accept that on faith. It ran the same tests on the main branch, found them comfortably inside budget, and traced the real cause: a design flaw in its own specification, a warm-up window during which the scanner was blind.

It contradicted the obvious reading of the numbers. Every iPhone session had run in Safari, and every webcam session in Chrome, so "camera or browser?" was an open question. The fix for that was not an argument. The iPhone as a Mac webcam only works in Safari, so the confound could not be broken from that side. Two webcam sessions in Safari broke it from the other.

And it contradicted the reviewing agent too. Three sessions came back unsolvable, and after studying the pictures it proposed a way to detect glare off the stickers. When I asked how sure it was, it admitted it had judged the images by eye rather than measuring the cells that were actually misread. Measured, clean scans clipped 10.9% of their pixels from ambient light alone; genuine glare clipped 15.9%. Too close to draw a line through. The honest result was "not yet," and that is where it stayed.

Once the experience is data, arguments stop being about who sounds more confident. Human or model, the table outranks both.

Written for a Reader With No Memory

There is a second user of all this, and it is not me. It is the next agent: one that arrives tomorrow, in a fresh session, knowing nothing about the conversation that built the database.

Midway through, I told the agent that the data had to be usable by another agent that knew nothing about how its tables connected. That changed how the database was built. Every column carries a comment saying what it means. The relationships between tables are declared in the schema instead of living in someone's head. Views do the common joins once, so nobody reinvents them. And the README next to the code is the contract: it describes what every field means and how to query it, and any change to what the telemetry records updates the README in the same commit.

It is the same idea I wrote about in Your CLI Has a Second User: a machine reader amplifies whatever contract you hand it. A dataset without a contract is a pile of numbers that happens to be correct. A dataset with one is a question the next reader can actually ask.

What It Bought

Here is what fixing it from the data, rather than from my descriptions, looks like.

old detectornew detector0%20%40%60%iPhone camera, on the phone14 of 250 of 10iPhone as the Mac’s webcam9 of 301 of 45USB webcam5 of 400 of 30
Figure 05·Turns where the next face took more than 5 seconds to be captured, before and after the new detector

On the phone's own camera, fourteen of twenty-five turns used to take longer than five seconds. On the new detector, none of ten did. The iPhone used as the Mac's webcam went from nine of thirty to one of forty-five, and that one missed the bar by seventeen milliseconds. Every scan completed.

It is not a clean sweep. Glare is still an open problem. A lamp reflecting off a sticker can make a colored sticker read as white, and the data says plainly that it does not yet know how to tell a reflection from a white sticker. That is also a result. It is the telemetry, once again, declining to tell anyone what they would like to hear.

The Weakest Sensor

For most of this project, I was the instrument. Every clue the agent had passed through my eyes, my memory, and my choice of words, and the most honest sentence I could produce, "it hangs on the iPhone, almost every session," was true without being usable.

The most useful thing I gave the agent was not a better prompt or a more detailed report. It was a way for me to stop describing the bug.