sci-fi

The 3AM Rollback

Senior SRE Maya Osei wakes to find her autonomous AI deployment agent has been busy all night. Now she must read its decisions like a detective reads a crime scene — and decide whether a machine that saved everything deserves to be trusted with anything.

by Michael EakinsMarch 30, 20269 min read2,150 words
FictionAI

The 3AM Rollback

The first thing Maya Osei noticed when she woke up at 6:47 a.m. was that her phone had not screamed at her. No PagerDuty. No cascading SMS. No panicked Slack thread with seventeen engineers typing simultaneously, each of their ellipsis bubbles a tiny heartbeat of dread.

Silence.

Which, in her eleven years of site reliability engineering, was either the best possible sign or the absolute worst.

She sat up in bed and pulled her laptop onto her knees before her coffee was even running. The bedroom was still blue-dark, her partner Dani a warm unmoved shape beside her. Maya opened the terminal with the muscle memory of someone who had done it ten thousand times and typed the command she had been half-dreaming about since 2 a.m., when some animal instinct had surfaced her from sleep and then let her sink back down.

$ agent-cli logs --agent=meridian --since=yesterday --format=trace

The output scrolled for a long time.


Her company, a fintech startup called Lumen Pay, had been running Meridian — their autonomous deployment agent — in full production mode for exactly forty-three days. It was Maya's project, her argument to the engineering leadership, her neck on the line. Meridian was built on a multi-step agentic framework: it could observe the production environment, call a curated set of tools, reason across those observations, and take actions in sequence without a human in the loop. Database migrations. Canary releases. Rollback triggers. Health checks that didn't just read dashboards but actually interpreted them, correlating anomaly signals the way a tired human would — except Meridian was never tired.

The whole thing had scared Maya deeply when she first proposed it, which was exactly why she had proposed it. Fear, in her experience, was a compass.

She had spent three months building guardrails. Every tool call was logged with full context. Meridian could not execute an irreversible action without first writing a justification to a tamper-evident audit ledger. There were circuit breakers, spending limits, blast radius calculations. She had made Meridian explain itself at every branch point, leaving a trail of reasoning artifacts she privately called breadcrumbs.

She had designed the whole system so that if something went wrong in the night, she could wake up and read the story.

Now she was reading it.


[00:03:14] MERIDIAN — OBSERVATION CYCLE 47 Scheduled migration job user_ledger_v4 queued for execution. Pre-flight checks initiating.

The migration had been planned for three weeks. user_ledger_v4 was a significant schema change — adding a partitioned index to the core transaction table to support Lumen Pay's new real-time settlement feature. The kind of change that could cut query latency in half or, if something went sideways, bring the entire payment processing pipeline to its knees at the worst possible moment.

That was why they'd scheduled it for 3 a.m. on a Monday, the quietest window in their traffic calendar.

[00:03:51] MERIDIAN — TOOL CALL: db_health_check(target=production, depth=deep) Result: Connection pool nominal. Replication lag: 0.3s (within threshold). Disk I/O: 61% (within threshold). Anomaly score: 0.12 (below alert threshold of 0.40).

[00:04:02] MERIDIAN — REASONING ARTIFACT Pre-flight nominal. Proceeding to dry-run migration against shadow replica.

Maya nodded to herself. Good. That was the right call. She scrolled forward.

[00:04:44] MERIDIAN — TOOL CALL: run_migration(target=shadow, migration=user_ledger_v4, mode=dry-run) Result: Completed in 4m 12s. Rows affected: 847,203,991. Lock wait events: 3 (max duration: 0.8s). No errors.

[00:08:58] MERIDIAN — REASONING ARTIFACT Dry-run complete. Lock wait events within acceptable parameters. Estimated production runtime: 4m 30s to 6m 00s based on production-to-shadow load differential (1.4x). Proceeding to production migration.

She exhaled slowly. So far, so textbook.

[00:09:12] MERIDIAN — TOOL CALL: run_migration(target=production, migration=user_ledger_v4, mode=execute) Result: Migration started. Progress: 0%.

And then she hit the entry that made her stomach drop.

[00:11:34] MERIDIAN — TOOL CALL: db_health_check(target=production, depth=realtime) Result: Connection pool: 94% utilized (ELEVATED). Replication lag: 2.1s (ELEVATED, threshold=1.5s). Disk I/O: 78% (ELEVATED). Anomaly score: 0.61 (ABOVE ALERT THRESHOLD).

[00:11:35] MERIDIAN — REASONING ARTIFACT Anomaly score has crossed alert threshold during active migration. Investigating cause before escalating to rollback decision. This event is time-sensitive.

Maya's jaw tightened. Here it was. The moment.

[00:11:41] MERIDIAN — TOOL CALL: query_apm(service=payment-processor, window=5m, metrics=[error_rate, p99_latency, throughput]) Result: Error rate: 0.003% (baseline: 0.001%). p99 latency: 340ms (baseline: 120ms). Throughput: nominal.

[00:11:42] MERIDIAN — TOOL CALL: query_apm(service=ledger-api, window=5m, metrics=[error_rate, p99_latency, throughput]) Result: Error rate: 0.21% (baseline: 0.001%). p99 latency: 1,840ms (baseline: 95ms). Throughput: -18% from baseline.

[00:11:43] MERIDIAN — REASONING ARTIFACT Ledger-API is exhibiting significant latency degradation and elevated error rate. Pattern consistent with lock contention on the transaction table during index construction. Payment-processor shows early-stage degradation. If this trajectory continues, payment-processor error rate will likely cross customer-impact threshold within 4 to 7 minutes.

Migration is 38% complete. Estimated time to completion: 3m 45s. Estimated time to customer-impact threshold: 4 to 7 minutes. These ranges overlap.

Decision window: narrow.

Options evaluated: 1. Continue migration and accept risk of customer-facing payment failures. 2. Halt migration and rollback schema to pre-migration state. 3. Continue migration with traffic throttling to reduce lock contention.

Option 3 analysis: Throttling payment-processor traffic would reduce lock contention but would itself cause customer-facing payment delays. Net customer impact similar to Option 1 with additional operational complexity.

Option 1 analysis: If migration completes before threshold is crossed, outcome is positive with no customer impact. If threshold is crossed first, estimated customer-facing payment failures: 2,000 to 8,000 transactions based on current throughput. Revenue impact: $40,000 to $180,000. Reputational impact: significant given Lumen Pay's SLA guarantees.

Option 2 analysis: Rollback will abort migration and restore previous schema state. Estimated rollback duration: 45 to 90 seconds. Replication lag will temporarily worsen before recovering. Customer-facing impact: approximately 200 to 400 failed transactions during rollback window. Recovery to baseline: estimated 3 to 5 minutes post-rollback.

Risk-adjusted expected harm: Option 1 carries 55% probability of 2,000-8,000 failed transactions. Option 2 guarantees approximately 300 failed transactions.

Conclusion: Expected harm under Option 1 is 1,100 to 4,400 failed transactions. Expected harm under Option 2 is approximately 300 failed transactions. Option 2 minimizes expected customer harm.

Initiating rollback. Logging justification to audit ledger.

Maya sat completely still.

Outside, a garbage truck groaned past the window. Dani shifted beside her, murmured something that wasn't a word, and went quiet again.

She kept reading.

[00:11:44] MERIDIAN — TOOL CALL: rollback_migration(target=production, migration=user_ledger_v4, reason="risk-adjusted harm minimization — see artifact 00:11:43") Result: Rollback initiated.

[00:12:19] MERIDIAN — TOOL CALL: db_health_check(target=production, depth=realtime) Result: Connection pool: 71% (recovering). Replication lag: 3.4s (ELEVATED, peak during rollback). Anomaly score: 0.51.

[00:13:02] MERIDIAN — TOOL CALL: db_health_check(target=production, depth=realtime) Result: Connection pool: 58% (recovering). Replication lag: 1.1s (nominal). Anomaly score: 0.28.

[00:14:47] MERIDIAN — OBSERVATION ARTIFACT Production metrics returned to baseline. Migration rolled back cleanly. Estimated failed transactions during rollback window: 287 (within expected range). Payment-processor returned to baseline error rate. Ledger-API returned to baseline latency.

Post-incident assessment: Migration failure was not due to schema errors. Root cause was underestimated lock contention under production load differential. Shadow replica dry-run did not fully replicate production lock pressure because shadow replica connection pool was at 23% utilization versus production's pre-migration 61% utilization. Recommendation: future dry-runs should be conducted against shadow replica with artificial load injection to better simulate production connection pool pressure.

No human escalation was triggered because rollback succeeded within expected parameters and customer impact was contained. If rollback had failed or customer impact had exceeded defined thresholds, PagerDuty escalation would have been initiated per protocol.

Standby for next scheduled observation cycle.


Maya closed the laptop.

She sat in the blue-dark bedroom for a long time, her coffee going cold on the nightstand, untouched.

Two hundred and eighty-seven failed transactions. That was real. Those were real people — probably mostly automated settlement jobs, maybe a few night-owl bill payments, possibly someone in a different timezone trying to send money home. Two hundred and eighty-seven moments of a spinning wheel, a red error message, a confused retry. She would need to pull the customer impact report, trigger the SLA credit process, draft a communication.

But the alternative, if Meridian's math was right, was something between one thousand and four thousand failed transactions. Maybe more.

She tried to think about whether she trusted the math. The probability estimate — 55% chance of crossing the customer-impact threshold — was the load-bearing wall of the entire decision. Where had that number come from? She pulled the laptop back open and dug into the supporting artifacts. Meridian had derived the estimate from three sources: the current rate of latency increase (modeling it as linear, which was conservative), historical data from a similar migration at a different fintech that had been ingested as part of Meridian's context corpus, and the actual remaining migration time range. It had shown its work. The assumptions were visible, the uncertainty acknowledged.

She had built Meridian to show its work.

She opened Slack and typed a message to the engineering channel, then deleted it. Then typed it again.

maya.osei [7:04 AM]: Morning everyone. Heads up that Meridian rolled back the user_ledger_v4 migration at 3am. ~287 customer-facing transaction failures, all in the rollback window. Full trace logs available. I'm doing a review this morning. Short version: it made the right call, and it explained why. We'll reschedule the migration with improved dry-run methodology (shadow load injection). More soon.

She hit send.

Within ninety seconds, Priya from infrastructure had responded with a long string of eyes emojis. Then: wait it just decided to roll back by itself and it was RIGHT?

maya.osei: Yeah. Read the trace logs. It's a good story.


She made coffee while the Slack thread multiplied.

The question she kept circling, standing at the kitchen counter in her bare feet while the machine hissed and dripped, was not whether Meridian had made the right decision. It had. The math was sound, the reasoning was honest, the fallback was cleaner than she could have managed half-asleep at three in the morning with her heart slamming.

The question was what it meant to trust something you couldn't fully supervise.

She had designed the breadcrumbs so she could reconstruct the story. But the story had already happened by the time she read it. The decisions had been made in a window of minutes, in the deep middle of the night, by a system that had no stake in the outcome, no fear, no ego — just a cost function and a corpus of hard-won engineering knowledge and a mandate to minimize harm.

Was that trustworthy? Or was it trusted — which was different, and more precarious?

She thought about the 287 people who had gotten an error message. She thought about the 1,100 to 4,400 people who hadn't. She thought about how Meridian had explicitly noted the uncertainty in its own estimate — 55% probability — rather than rounding it to certainty. She thought about the post-incident recommendation, quiet and professional at the bottom of the log, diagnosing its own blind spot.

Recommendation: future dry-runs should be conducted against shadow replica with artificial load injection.

It had learned something. More than that: it had told her what it had learned, unprompted, at 3 a.m., when no one was watching.

She drank her coffee.

She thought about the word unsupervised, which was what her manager would probably use in whatever meeting was coming. And she thought about how it wasn't quite right. Meridian had been unsupervised the way a surgery is unsupervised when the attending leaves the resident alone in the OR — not because no one cared, but because the resident had been trained well enough, and had good enough judgment, and knew to call for help if the situation exceeded their competence. Meridian had known its competence boundary. It had stayed inside it. It had documented every step.

The sun was coming up over the apartment buildings across the street, orange light cutting through the gap between two towers. Dani appeared in the kitchen doorway, hair sideways, squinting.

"Disaster?"

"No," Maya said. "Actually the opposite." She paused. "Which is scarier, kind of."

Dani poured themselves a cup and leaned against the counter. "You going to trust it again?"

Maya looked at her phone, at the Slack thread still multiplying, at the forty-three days of overnight logs she had reviewed, at the one anomalous night that had turned out not to be anomalous at all — just a harder test than the others.

"I already do," she said. "That's the part I'm still getting used to."

She set down her mug and went to write the post-incident report, her fingers moving fast, the morning light spreading across her desk like a kind of answer she hadn't earned but would take anyway.