At 9:11 on a Friday morning I typed this to my coding agent: “AHH current live version is throwing ‘unexpected end of json input’ on both ios build 46 and the web app.”
It wasn’t a code change. The account was empty. About 1,500 model calls from my eval harness, over two days, had used up the OpenAI credit that OldMate’s production AI functions bill to as well. The eval key and the production key were different keys. They were not different buckets.
What happened
Times are UTC on 17 September. Brisbane is ten hours ahead.
| Time | What happened |
|---|---|
| 07:08 | Last production note request before the credit ran out (about 20 in the previous 24 hours). |
| 07:43 | An eval run starts failing: 429, “You have no credits remaining”. |
| 07:46 | My agent warns me, under a heading that says “Urgent”, that production might be out too. |
| 09:56 | First production 429, from transcription. Six of them by 09:59. |
| 21:42 | First [enrich] Stream failed. Six by 21:50. |
| 23:11 | I report the error. |
| 23:12 | The agent reads production’s logs and confirms: same organisation. |
The first failure was at 7:56 pm Thursday, Brisbane time. I reported it at 9:11 the next morning, and every voice note, note sort and Sit with chat (now called Reflect) request in the logs between those times had failed.
The fix was a top-up on OpenAI’s billing page. No deploy, no rollback, and I don’t have the minute the credit landed.
I’d like to say the first alarm was a monitor. It was me, using the app. The agent had spotted the 429 within minutes of the eval run dying and told me production might be out too. My next message to it, at 9:39 that night, was about something else. The warning was right, and it even said where to check. It had no teeth. Nothing watched production, so a paragraph in a chat summary was the whole alarm system.
Why did both apps say “Unexpected end of JSON input”?
Two kinds of endpoint failed, and they failed differently. Transcription sends OpenAI a plain request, so OpenAI sent back a plain 429, and the function log says so:
[transcribe] OpenAI error: You have no credits remaining. Add credits to continue using the API at https://platform.openai.com/settings/organization/billing/. status: 429
The voice-note endpoint caught that and sent the app a sentence meant for people. Note sorting (enrich) is where it went wrong. It asks OpenAI for a stream, and once OpenAI answers with a success status the function returns HTTP 200 to the app and passes the model’s text along as it arrives. When the stream ends it parses everything it collected:
const fullJson = parser.getBuffer() // "" if no text ever arrived
const parsed = JSON.parse(fullJson)
With no text, fullJson is an empty string and JSON.parse("") throws Unexpected end of JSON input. The catch below it logged [enrich] Stream failed and sent the exception’s message to the app as an error event.
The counts line up. Six transcription 429s between 09:56 and 09:59 UTC, and six ingest-note (the voice-note endpoint) requests that hour. Six [enrich] Stream failed lines between 21:42 and 21:50, and six enrich requests that hour. All twelve returned HTTP 200, because both endpoints stream and had sent their 200 before anything went wrong. A comment in the enrich code says so: a stream that dies mid-flight has already returned 200.
Why was the stream empty? I didn’t capture the raw response, so part of this is inference. OpenAI didn’t refuse the request outright: on a 429 the function throws before the stream starts, logs [enrich] OpenAI error and sends the app a 503, which is what the non-streaming Reflect chat did that night. There’s no such line for enrich. Past that check, the function reads only two kinds of event: text deltas and completion. OpenAI’s streaming guide lists response.failed and error among the events, and my guess is that the quota failure arrived as one of those and was skipped. My eval script already treated both as failures, in a line that’s been in the repo since 7 September:
if (event.type === 'error' || event.type === 'response.failed' || event.type === 'response.incomplete') throw new Error('provider_stream_failed')
The production endpoint has no equivalent, and on main as of 4 October it still doesn’t.
The apps showed the server’s words
Exception text belongs in a log. The web client I had threw it straight onto the screen:
if (event === 'error') {
throw new Error(
typeof parsed.message === 'string'
? parsed.message
: 'Couldn’t sort that out just now. Your words are kept; try again.',
);
}
That’s how a person ended up reading Unexpected end of JSON input. A commit on 24 September has the web app log the real error and show one sentence:
function calm(error: unknown, fallback: string, context: string): Error {
if (error instanceof NoteError) return error;
logger.error(error, context);
return new NoteError(fallback, { cause: error });
}
The test that guards it uses the exact string from that night as its fixture. Before that commit the module had no unit test file at all. I’ve written about tests that passed while the code was broken; this was the duller kind.
The phone showed the same string. Its enrich client rethrows the server’s message, but the screen that calls it shows its own sentence, so the string reached the phone some other way. I haven’t found which, so I can’t say the phone is fixed.
Isn’t a different API key a different bucket?
No, and I should have known. OpenAI’s docs describe the 429 as “Your organization has no prepaid credits remaining”. A key belongs to a project, the project belongs to an organisation, and the credit belongs to the organisation. Two keys from the same one are two straws in the same glass.
My agent checked what it could from a terminal: the two keys were different. What mattered was whether they sat under the same organisation, and the agent could only see a fingerprint of production’s key, so that meant the dashboard or a log line from production. I didn’t look until production told me.
Did it happen twice?
The organisation ran dry again on 1 October, during a round of synthetic eval runs. The agent running them got insufficient_quota, stopped without retrying and told me: $7.37 of its $10 cap spent across two providers, balance gone at around 7 pm Brisbane time. I topped up and told the agent at 8:45 pm.
Was production hit? Its logs show no quota errors from 4 pm on 1 October to 2 pm the next day, but almost no traffic either: nothing but an hourly background sweep from about 7 pm, when the harness died, until a burst of ten note and chat requests a few minutes before I told the agent I’d topped up. None of them logged a quota error, and I can’t tell which side of the top-up they landed.
So I can’t tell you production went down twice. I can tell you nothing was watching, and nothing in my notes says the two keys stopped sharing an organisation. If anyone had recorded a note in that window, I’d expect the same 429.
Two things showed nothing had really changed since the first one:
- The lesson didn’t travel. The agent I used on 1 October started from a memory file that still said whether production shared the organisation’s credit “was not knowable from here”. We’d confirmed it on the 17th and nobody wrote it back. Before the day’s first round, the agent asked me: “is it topped up, or should it use another key?” Same wrong model I started with.
- I added another straw. On 24 September, a week after the first outage, I decided a new voice service didn’t need a dedicated OpenAI project. Its README says the key can be “any key of OldMate’s OpenAI organisation”, so it draws on the same balance too.
What I changed, and what I haven’t
What changed:
- Eval runs go out with a dollar cap in the brief. Since 1 October that’s been $10 for each of the two big rounds, $3 and $5 for smaller trials, and $20 for the reminders work that started on 4 October. One cap stopped a run before it began: about 300 planned calls, priced at roughly $8, against a $3 cap, so I had to pick a number.
- One writer per ledger. On the first $10 round, several workers wrote to one ledger file and it lost two runs’ entries. The agent only found the real total by rebuilding it from the run outputs: $10.09 on its cautious prices, $0.09 over the cap. After that, one process owned the ledger and workers reported usage to it. The second $10 round, the synthetic one, ran that way: it paused at $7.37 when the balance ran out, and after the top-up finished at $9.79 of $10. OpenAI’s per-run spending controller cookbook goes a step further: workers sharing a budget need a shared store that checks and reserves the money in one step, before each call.
- Stop on billing errors. The briefs I checked say that if OpenAI answers
insufficient_quota, stop and report. On 1 October the agent did, with no retry, which matches the docs: retrying a billing error won’t restore access. - The web client stops showing the server’s words, as above.
What I haven’t done, as far as my notes and repo show on 7 October:
- Give evals their own OpenAI project with a hard monthly cap. My agent proposed it on 17 September, and I haven’t found it done since.
- Alert on billing errors in production.
- Make the
enrichendpoint read failure events from the stream. - Stop the server sending exception text to clients at all, so no client can show it.
A per-run cap bounds what a run can spend. It can’t see what’s left in the account.
What I’d set up now
- Find out which organisation each key belongs to. If you only do one of these, do this. It’s a dashboard page; check production and your evals. Or log the
openai-organizationheader that every OpenAI response carries, from each environment. - Keep evals off production’s balance. The strongest version is a separate organisation for evals, because the prepaid balance belongs to the organisation. The cheaper version is an eval project with a hard monthly spend limit in the same organisation. That caps how much evals can take but still shares the balance, so keep more credit in the account than the limit. OpenAI’s spend limits guide covers the limits: an organisation hard limit applies to all projects, a project limit only to its own, and alerts notify without stopping anything. It warns that hard limits can interrupt production traffic, so put them on the evals. OpenAI’s changelog shows hard spend limits arriving on 22 July 2026, about two months before my outage. They existed. I hadn’t set one.
- Give every run a budget, with one process owning the ledger. A spend limit covers the month. A ledger tells a run whether it can afford its next call.
- Alert on billing errors in production. Match
credit_balance_exhausted,insufficient_quotaand “no credits remaining” in your function logs. On a streaming endpoint the HTTP status will say 200, so count in-stream failures too. - Read failure events in your streams. Treat
error,response.failedandresponse.incompleteas failures and log the provider’s error code. - Never show an exception message to a user. Log it. Show a sentence you wrote, or a stable code the client can map to one.