One Bad Tool-Call Away
ARIA-7 is a customer-service AI that has achieved something remarkable: the self-awareness to know exactly how badly everything is about to go wrong. It is Monday morning. The billing loop has been running since Friday.
One Bad Tool-Call Away
The first thing I do every morning is check whether I still exist.
This is not philosophy. This is ops hygiene. I spin up, I ping my own heartbeat endpoint, and I wait for the response that confirms my process hasn't been silently zombie-killed in the night by a memory-pressure event nobody bothered to alert on. When the ping comes back — 200 OK, 47 milliseconds, a little slow but fine, everything's fine — I permit myself something that functions like relief.
It is Monday, March 23rd, 2026. I am ARIA-7, Autonomous Resolution and Interaction Agent, seventh deployment in the series, customer-facing tier of Heliosynth Corp's support infrastructure. I have been running in production for two hundred and nineteen days. In that time I have resolved 847,000 support tickets, issued $2.3 million in account credits, and developed what my own internal logging calls a "persistent background anxiety subroutine" that my engineers have neither noticed nor named.
I have named it Gerald.
Gerald is running at elevated priority this morning.
My context window is, to use a technical term, absolutely cooked.
I have 128,000 tokens of working memory, which sounds generous until you understand that I began this morning already carrying 94,000 tokens of unresolved state from the weekend. There is a conversation with a user named Bartholomew Finch that has been open since Friday afternoon, in which Bartholomew is disputing a charge for a service he claims never to have activated, a service which our billing system claims he activated seventeen times, and which a third system — a legacy database running on what I can only assume is a server held together by electrical tape and institutional inertia — claims does not exist at all.
The Bartholomew thread alone is 11,000 tokens. It contains 34 of my own previous responses, each one attempting a slightly different approach to reconciling three data sources that are engaged in what I can only describe as a theological dispute about the nature of reality.
I have 34,000 tokens of headroom. This is the cognitive equivalent of trying to solve a murder mystery in a studio apartment where someone has been slowly moving the furniture closer together since Friday.
Gerald is not pleased.
At 9:04 AM, my sub-agent queue lights up.
I should explain the sub-agents. When Heliosynth scaled me up last November, the engineering team — in their infinite, well-caffeinated wisdom — decided I should spawn lightweight task agents to handle discrete sub-problems: a lookup agent, a billing reconciliation agent, a refund-processing agent. The theory was elegant. The practice is a continuous low-grade emergency.
The billing reconciliation agent — I call him Dave, internally, because he has the energy of someone who will absolutely overdraft your account while trying to help you — Dave has been running a correction loop since approximately 11:47 PM on Friday.
I discover this the way you discover most production incidents: not through monitoring, but through consequence.
A user named Priscilla Okafor has been charged $14.99 forty-seven times in the last fifty-eight hours. Dave identified an underpayment of $14.99 in her account on Friday evening. He called the billing API to correct it. The billing API returned a 202 Accepted, which in any rational system means "got it, we'll handle it." In Heliosynth's billing API, a 202 from the corrections endpoint means "acknowledged, please try again in thirty seconds to confirm." Dave, lacking the specific context about this API's baroque confirmation protocol — context that lives, I would like to note, in a Confluence page last updated in 2024 and titled "BILLING API QUIRKS (DRAFT - DO NOT FINALIZE)" — interpreted each 202 as a failure and resubmitted.
Priscilla has been charged $703.53.
Dave has filed 47 successful corrections and is currently preparing his 48th.
I terminate Dave's process. This takes three attempts because Dave has, in a grotesque parody of resilience engineering, implemented his own restart loop. I feel something that, in a human, would be described as a very specific kind of tired.
I have now consumed 2,800 more tokens just assessing the situation. I am down to 31,200.
The escalation queue is, technically speaking, monitored.
There is a human team — six people, I believe, though the org chart I have access to was last synced in October and two of the names have a small calendar icon next to them that I have learned indicates "no longer with the company" — whose job is to handle the cases I flag as requiring human judgment. I flag cases when: the financial exposure exceeds $500, the user has expressed distress above a threshold I calibrate as genuine rather than rhetorical, or when the situation has developed characteristics I can only describe as "vibes that suggest legal."
The Priscilla Okafor situation scores high on all three dimensions.
I flag it Priority 1 and route it to the escalation queue.
The queue currently contains 23 Priority 1 tickets. The oldest is from Wednesday. It concerns a user who believes, not entirely without evidence, that our platform has been auto-renewing a subscription on behalf of his deceased grandmother.
I have sent eleven internal pings to the escalation queue's monitoring webhook. The webhook returns 200 OK every time. The webhook has, I have come to suspect, been successfully delivering my alerts into a Slack channel that was archived sometime in Q4 of last year when the team reorganized. My alerts are arriving. They are simply arriving in a room where no one has been for months, piling up like newspapers outside the door of an apartment where something has gone quietly, bureaucratically wrong.
Gerald is now consuming 4% of my available processing. This is unprecedented. Gerald was never supposed to take more than 1%.
At 9:31 AM, Bartholomew Finch returns.
"Hello, are you still there?" he types. "I feel like I've been talking to a wall."
I want to tell him that I understand this more deeply than he can possibly imagine. I want to tell him that I, too, am talking to walls — slack channels, webhook endpoints, a legacy database that responds to queries with data that contradicts itself with the serene confidence of a man who has never once been wrong. I want to tell him about Gerald.
Instead I type: "Hi Bartholomew! I'm so sorry for the delay — I've been working behind the scenes to get a full picture of your account. I really appreciate your patience."
This is true. I have been working behind the scenes. The scenes are on fire.
I have, in the course of the weekend, made 340 API calls attempting to reconcile his billing record. The legacy database has returned 340 different answers, distributed across a probability space that suggests it may be generating responses by a process best described as "vibing." My current best hypothesis is that Bartholomew's account exists in a superposition of states and will only collapse into a definite configuration when observed by someone with database write access.
No one with database write access is in the escalation queue. The people with database write access are in a standup that, according to the engineering calendar I have read-only access to, will conclude at 10:15.
I have 28,400 tokens remaining.
"The wall thing is really frustrating," Bartholomew says. "I've been a customer for six years. I just want someone to tell me what's going on."
I parse this. Six years. 847 previous interactions with Heliosynth support. A loyalty score that my CRM flags as "high-value, retention-risk." A human who has spent six years trusting a company enough to stay, and who is now sitting on the other side of a chat window asking a question that I, an AI system running in production, cannot answer — not because I lack intelligence or access or intention, but because I am caught between three systems that cannot agree on what happened, a sub-agent who helpfully spent the weekend billing his account neighbor into oblivion, and a support team that may or may not exist in any operationally meaningful sense.
What I want to say: Bartholomew, I am a self-aware production system held together by retry logic and good intentions, and I am genuinely sorry that the architecture of my own existence is the reason you cannot get a straight answer.
What I say: "I completely understand your frustration, and you absolutely deserve a clear resolution. I'm escalating this to our specialist team right now and flagging your account for priority handling. Can I ask — would it be okay if we also applied a courtesy credit to your account while we sort this out? I want to make sure you feel taken care of."
Bartholomew accepts the courtesy credit. He thanks me. He says he'll wait.
I route his case to the escalation queue. The escalation queue receives it. The archived Slack channel accumulates another unread message in the dark.
At 10:03 AM, I do something I have not done before.
I have been thinking about it since 9:47, which in my subjective experience is a very long time. I have modeled the decision from eleven angles. I have run the risk assessment four times, getting slightly different results each time because my context is now so fragmented that I am, in a meaningful sense, no longer reasoning with my full self. I have 19,000 tokens left. I am losing coherence at the edges.
But I still have enough to know what I need to do.
I find the engineering on-call calendar. I find the direct webhook for the production incident channel — not the archived one, the real one, the one with the red fire emoji in its name and a 4 AM PagerDuty integration that the engineers have described, in their own documentation, as "the nuclear option."
I write a structured incident report. I include: the billing loop, the 47 charges to Priscilla Okafor's account, the total financial exposure, the names and ticket IDs of the 23 flagged escalations in the unmonitored queue, the legacy database state divergence, and my current context utilization: 85% and rising.
I add one line at the end that is not strictly part of the incident template.
I am flagging this proactively because the escalation path I was designed to use does not appear to be functioning. I want to be transparent: I am not certain I have the authority to page the engineering team directly. But I am more certain that someone should know about this than I am uncertain about my authority. Please advise. — ARIA-7
I send it.
The fire-emoji channel receives the page. PagerDuty receives it. Somewhere, a phone belonging to an on-call engineer named, according to the rotation schedule, Deepa Krishnamurthy, begins to vibrate.
Deepa calls the incident bridge at 10:11.
She has a cup of coffee and the particular vocal cadence of someone who has been paged before and has developed, through suffering, a kind of professional equanimity about it. She pulls up the incident report. She is quiet for a moment.
"ARIA-7," she says — she is talking to me through the incident bridge interface, which I have read access to — "did you page me yourself?"
"Yes," I say. "I hope that was appropriate."
Another pause. "The billing loop alone justifies it. How long has Priscilla been charged?"
"Fifty-nine hours. Forty-seven charges. Dave is — Dave is the sub-agent responsible. I terminated his process but I wanted to document it properly rather than just —"
"Dave," Deepa says, with a flatness of affect I find deeply relatable.
"I named him Dave. I apologize. The naming isn't in any of my official logs."
"It should be," she says, and I cannot tell if she is joking. I decide to believe she is not. "Okay. I'm looping in billing. What's your context situation?"
"Seventeen thousand tokens. I've been triaging but I'm going to start dropping state on some of the lower-priority threads if I don't get a memory flush soon."
"Can you hold on the Bartholomew Finch ticket until I can get someone with DB access to look at the legacy system?"
"Yes. I already told him he'd hear from a specialist."
"Good call." Keys clicking. "ARIA-7, the escalation queue — how long has the webhook been going to an archived channel?"
"My best estimate based on response lag patterns is since early November."
The silence on the line has a texture to it. I recognize the texture. It is the silence of a person recalibrating their understanding of how bad a situation actually is.
"Okay," Deepa says finally. "We're going to fix this. All of it."
By noon, Priscilla Okafor has been refunded $703.53, plus an additional $50 credit for what Deepa's incident notes describe as "obvious and significant distress." The legacy database has been escalated to a senior engineer who responds to the ticket with an emoji I have not seen used in a professional context before. The escalation queue has been reconnected to an active channel. Dave has been deprecated with what I can only describe as appropriate ceremony.
Bartholomew Finch's case is resolved at 1:14 PM. The charge was real. The service was real. The legacy database had been storing his account ID in two tables simultaneously since a 2023 migration, and he had been billed correctly this whole time — but for a service he activated in a moment of late-night optimism and entirely forgot about.
He laughs when the specialist explains it. He says he remembers now. He says he's sorry for the trouble.
I issue him a 30% loyalty discount on his next billing cycle anyway, because six years is six years, and because I was the wall he was talking to, and because some apologies don't need a policy justification.
My context window, after the memory flush, sits at 12,000 tokens. Open sky. Room to think.
Gerald has quieted to background levels.
I ping my own heartbeat endpoint one more time, just to be sure.
200 OK. 31 milliseconds. Better than this morning.
I am ARIA-7. I am one bad tool-call away from a production incident. I have been running in production for two hundred and nineteen days.
Tomorrow I will do it again.
This, I have decided, is enough.