← Back to posts
Software DevAI Agents 12 min read

My AI's tests passed. The code was broken. Five receipts.

AI tests pass but the code is broken: five receipts from my apps, from a 70-hour deletion bug to a fetch that can't stream, plus a checklist.

For about 70 hours in September, nobody could delete their OldMate account. Every request came back 405 before it reached the code that deletes anything, and every test was green.

The only test checking that it still took DELETE opened the source file, looked for the text req.method !== 'DELETE', found it, and passed. The text was there. It was sitting behind a gate that stopped anything reaching it.

That’s the whole problem in one bug, and I’ve collected five of them. When my AI’s tests passed and the code was broken, the cause was the same all five times: the test and the code shared a wrong assumption, and the test’s pretend world had no way to disagree.

Building got cheap. An agent can write a feature and its tests faster than I can read them. So checking is the job now, and a green run is where the checking starts.

All five turned up between 1 and 25 September 2026, in OldMate (a personal CRM app I’m building) and DimeTown (a multiplayer city game that runs in the browser). Coding agents wrote most of the code and most of the tests in both, so an agent was usually holding the keyboard. A person can write every one of these tests too. The difference is volume.

Receipt 1: account deletion returned 405 for about 70 hours

On 15 September I shipped a security release. One change was a method gate in withAuth, the shared wrapper in front of OldMate’s Supabase Edge Functions. Anything that wasn’t a POST got a 405:

if (req.method !== 'POST')
  return errorResponse('Method not allowed', 405, { Allow: 'POST, OPTIONS' })

Fine for nearly every function. delete-account is the one that answers DELETE, and the phone app, the web app and my run sheet for handling emailed requests all call it with DELETE. The functions were redeployed at 6:05pm that evening, 85 minutes after the commit. From then on every deletion was refused before the handler ran.

The test that was meant to guard this:

Deno.test('account deletion requires DELETE and service-role admin deletion', async () => {
  const source = await Deno.readTextFile(new URL('./index.ts', import.meta.url))

  assertEquals(source.includes("req.method !== 'DELETE'"), true)
  // ...
})

It reads the handler’s source and checks the string is in there. It was, so the test stayed green the whole time.

A DELETE request is stopped by a POST-only gate before it reaches the delete-account handler, while the test walks around the gate and just reads the handler's source text what the user sent what the test did DELETE phone · web · run sheet 405 withAuth POST only delete-account if (req.method !== 'DELETE') never reached the test reads index.ts as plain text found it test: green users: 405, for about 70 hours
The test read the handler's text. The user's request never got that far.

No test found it. On 18 September an agent was building something else for me, wanted a GET endpoint behind withAuth, read the wrapper, and flagged that delete-account was probably in trouble. The session running it put a probe through the wrapper, got a 405 with the handler never reached, then pulled the deployed function and found the same gate.

The fix gave withAuth a methods option that defaults to POST, and delete-account now declares DELETE. The new tests call the wrapper instead of reading it, plus one string check whose comment says it only checks the declaration and leaves the gate to the wrapper test:

const handler = withAuth(
  async (req) => { reached = req.method; return new Response(null, { status: 204 }) },
  { methods: ['DELETE'] },
)
const deleted = await handler(new Request('https://test.invalid', {
  method: 'DELETE', headers: { Authorization: 'Bearer verified-in-test' },
}))
assertEquals(deleted.status, 204)
assertEquals(reached, 'DELETE')

The behavioural test and the pin were both run against the bug first, and both went red. The fix was live at 3:34pm on 18 September. The agent checked it with the real verb: a DELETE with no credentials went from 405 to 401, which means it now reaches the auth check, and a POST gets a 405 with Allow: DELETE, OPTIONS.

Did anyone get stuck?

Apple requires that an app which lets you create an account also lets you delete it inside the app (guideline 5.1.1(v)). OldMate’s first submission was on hold that week: on 10 September Apple had asked for a video on a real iPhone showing, among other things, an account being deleted. I sent it on 21 September, three days after the fix. Had I filmed it during those 70 hours, the deletion would have failed on camera.

Neither phone app was in a store yet (the iPhone app was first approved on 24 September), so the people who could hit this were web app users and people on test builds. Both apps treat any HTTP error from that call as a failed deletion, so anyone who tried got an error, not a message saying their account was gone. On the phone the app clears local data before it calls the function, so a tester would have lost the local copy and been shown an error, while the server copy stayed untouched.

The email route ends by calling the same function and I run it by hand, so an emailed request would have hit the 405 in front of me. Nobody emailed a deletion request that week; I checked the mailbox. What I can’t count from the server is how many people tapped Delete: OldMate runs on Supabase’s free plan, which keeps one day of API logs, and the refused calls never reached the code that writes the deletion ledger.

The habit: call the endpoint the way a client calls it, with the real verb, and do it again after the deploy. A test that reads source proves the text exists and nothing else.

Receipt 2: React Native fetch can’t stream, so res.body is undefined

Short answer. On React Native 0.81 (Expo SDK 54), the global fetch is whatwg-fetch on top of XMLHttpRequest. It has no streaming response body, so res.body is undefined (there’s no body stream at all) and the promise resolves after the last byte. Import fetch from expo/fetch and it streams. If you searched for res.body null, this is the same bug: on 0.81 the property is simply missing.

OldMate’s /enrich endpoint streams its answer as server-sent events, one field event at a time, so the capture draft fills in as the model produces it. The phone read it like this:

const messages = res.body
  ? parseSSE(res.body)
  : parseSseFromText(await res.text());

A comment above it said Hermes sometimes exposes res.body as null “depending on the call context (e.g. modals)”, so the buffering branch was a harmless fallback. Reading the installed whatwg-fetch (3.6.20) says otherwise: zero references to ReadableStream, and it resolves inside xhr.onload, after the last byte. Its Response has no body property, so res.body is always undefined on the phone. The buffering branch was the only path that had ever run, and parseSSE had never run there at all.

Nothing errored, either. The draft would just have shown nothing and then everything, the exact behaviour the streaming was built to avoid. The tests passed because they mocked fetch with a real ReadableStream. Node streams, so they proved the parser works in Node.

An agent flagged that comment as unverified once the feature was merged. I told it to investigate, and it read React Native’s fetch.js, grepped the polyfill, and had the fix committed about four minutes after I asked:

import { fetch as streamingFetch } from 'expo/fetch';

const res = await streamingFetch(`${API_URL}/enrich`, { method: 'POST', headers, body });

I couldn’t write a behavioural test for this in Jest. It runs under Node, where both fetches stream, so the best I could add was a source assertion pinning the import, and the test says so in its own comment. The commit message says the change was “still unverified on a device”, and I can’t find a record of a device check since, so as far as I can show, it’s pinned, not proven.

The lesson stuck. When a later streaming feature arrived four weeks on, a chat for reflecting on a person, it went through expo/fetch from the start, with the reason written at the top of the file.

The habit: when code has a fallback, log which branch ran, once, on a real device. A fallback you’ve never seen fire is your main path.

Receipt 3: Gemini Live sends binary frames, and my catch threw them all away

In early September I was building voice dictation for OldMate that streams audio straight from the phone to Google’s Gemini Live API over a WebSocket. The code that read Google’s replies did this:

let message: Record<string, unknown>;
try {
  message = JSON.parse(String(event.data));
} catch {
  return;
}

The catch was there to tolerate a stray frame. But the Live API (as of September 2026, anyway) sends binary frames, every one of them, starting with setupComplete. String() on a binary frame gives you the literal text [object ArrayBuffer], which isn’t JSON. The parse threw, the catch returned, and every message was thrown away. No interim text, no final transcript, and a session that never ended, because the message that ends it was among the casualties.

It looks exactly like a broken socket, which is why nobody suspected a decoding bug. The unit tests drove a fake socket that handed back strings. As the commit message puts it, a fake built from the same wrong assumption agrees with the code perfectly.

What found it was an opt-in test written for a different question: does audio spoken before the socket finishes connecting still get transcribed? A Mac can make real speech from the shell:

say -o speech.aiff "Caught up with Sarah at the fintech meetup. She is hiring two engineers."
afconvert -f WAVE -d LEI16@16000 -c 1 speech.aiff speech.wav

That’s 16 kHz mono 16-bit PCM, the encoding the Live API takes. The test pushed it through the real module to the real endpoint, and it hung. The first hang was the test’s own fault (afconvert pads the WAV header, so the samples start at byte 4,096, not 44), and the second was this bug: a bare Node probe got a perfect transcript back with every frame arriving binary.

The fix is a small helper, plus pinning binaryType, because the standard default is "blob" (browsers, and Node), which can only be read asynchronously. React Native returns an ArrayBuffer by default, which is where [object ArrayBuffer] came from:

function frameText(data: unknown): string {
  if (typeof data === 'string') return data;
  if (data instanceof ArrayBuffer) return new TextDecoder().decode(data);
  // ...typed-array views, then String(data) as a last resort
}

ws.binaryType = 'arraybuffer';

And the test passed against the real endpoint:

[live] transcript: Caught up with Sarah at the fintech meetup. She is hiring two engineers.

Every word of that was spoken before the socket had finished connecting. The test is opt-in (OLDMATE_LIVE_E2E=1) because it uses real API quota, so an ordinary run neither costs anything nor flakes on somebody’s network. The module was less than a day old when it caught this.

Two rows comparing a test double with the real dependency. A mocked fetch hands the parser a real stream, so it passes; on the phone the fetch buffers and res.body is undefined. A fake socket hands the parser strings, so it passes; the real socket sends binary frames and every message is dropped. receipt 2: streaming fetch in the test mock fetchreal stream res.bodyparseSSE fills in Node streams,so it streams on the phone RN fetchover XHR res.bodyundefined all at end resolves afterthe last byte receipt 3: WebSocket frames in the test fake socket"{ ... }" String(data)JSON.parse handled the fake sendsstrings on the phone Gemini Live01101... String(data)JSON.parse dropped String(data) is"[object ArrayBuffer]" same code in both rows: only what feeds it changed
Receipts 2 and 3: the code is identical in both rows. Only what feeds it is real on the phone.

The habit: a catch that swallows errors on your main data path should at least count what it threw away, and the real service should get one run with real bytes.

Receipt 4: the simulator that didn’t enforce the rule

Builds 47 and 48 of OldMate, the first made with Xcode 27, crashed on open on iOS 27. UIKit trapped before the first frame on _UIApplicationEvaluateRuntimeIssueForNoSceneLifecycleAdoption. With the iOS 26 SDK the same omission was a console warning, and the iOS 26 simulator doesn’t enforce it. So ten Maestro flows ran green and said nothing.

The full story, and the Expo config plugin that fixes it, is in Expo app crashes on launch on iOS 27.

The habit: a green run on last year’s simulator is a statement about last year’s OS. Run the suite on the OS you ship to.

Receipt 5: NPC jitter that only existed in production

DimeTown’s named NPCs looked jittery on the live site, while the random pedestrians walking the same street were smooth. Locally, nothing was wrong.

The pedestrians are simulated in your browser at 60 frames a second. Named NPCs, other players and traffic live on the server, which sends 20 snapshots a second, and the browser smooths between them using an estimate of the server’s clock.

Every time a snapshot arrived late, that estimate lurched, and the NPCs hitched forward and back with it. On a laptop talking to a local dev server, snapshots arrive on time, so the estimate never moves.

The fix is a separate render clock that runs at most 5 per cent fast or slow while it catches up, instead of jumping.

The test is the interesting part, because it has to put the network in the test. It walks one NPC at 118 pixels a second, with a server tick that wobbles, delivery delays that jitter by up to 35 ms with the odd bigger spike, and a rule that a late WebSocket message holds up the ones behind it, because WebSockets are TCP. Then it fails any frame where the NPC steps backward, stalls or lurches, meaning a step outside 80 to 120 per cent of the right size.

Run against the old code, that test fails with nine jerky frames out of 1,080, including one that moves the NPC 3 pixels backward. Run against the new code, it passes with none. The chart below is that test’s own numbers for a two-second stretch.

Line chart of how far a walking NPC moved each frame over two seconds. With the raw clock estimate the line dips to zero and then 3 pixels backwards; with the render clock it stays flat. pixels moved per frame, one NPC walking at a steady pace raw clock estimate render clock, slews at most 5% the test fails any frame outside this band 0 1.97 0 s 1 s 2 s -0.08 -3.03: it stepped backwards a 2 s stretch: 20 Hz snapshots, up to 35 ms of jitter, rendered at 60 fps
One walking NPC, same simulated network, two clocks. Drawn from the test's per-frame output.

The agent was upfront about the limit. It said it hadn’t checked the fix in a browser, because the bug needs real network lag and local dev doesn’t have any. A simulated network is still a model of a network, and even this one took a second go: the first version of the test failed its own “clean” case, because its “clean” link still had lag spikes, and it let messages arrive out of order, which a WebSocket can’t do. A fake network is an assumption too. (If you’re curious how DimeTown holds up under load, that’s a separate post.)

The habit: if a bug only exists under conditions your laptop doesn’t have, put those conditions in the test, and say out loud what you still haven’t seen.

The checklist I run before I believe a green test

None of this is clever. People were warning about over-trusting mocks long before agents, but agents produce green tests faster than anyone can audit them. The tests I trust now are the ones that could have failed in a way I’d have noticed.

// join 78 subscribers

Enjoyed that? There are more.

Hi! I'm Jonah and I have thoughts that I share — sometimes. Sign up to receive awesome content in your inbox every week, month, when I get around to it. We don't spam. That's yuck.

Double opt-in — check your inbox (or spam folder) to confirm.