Go-live day ends with applause. At 3am the agent loops on a broken tool, inference spend spikes, or retrieval returns empty and the model improvises policy answers. Who gets the page? What do they do first? Who can roll back?
If the answer is "we will figure it out," you are not in production. You are in an experiment with production data.
Operations is part of architecture
Runbooks, on-call rotations, escalation paths and rollback procedures belong in the build the same way authentication does.
Discovery asks:
- Expected support hours and regions
- Criticality class of the workflow
- Who operates — client ops, client IT, Incubics managed run, mix
- RTO/RPO expectations for the AI service and its dependencies
Ambiguity here becomes pager pain later.
Runbooks operators can follow
A runbook is not a wiki essay. It is steps:
- How to confirm the symptom — dashboards, logs, sample cases
- Known failure modes and fixes — refresh corpus, disable tool, route to manual queue
- Rollback — redeploy previous model/prompt version; expected time
- When to escalate to engineering
- Customer comms template if user-facing impact
Runbooks are tested in dress rehearsals before go-live, not written after the first incident.
Monitoring beyond uptime
Green health checks lie. Useful signals:
- Tool error rate and latency
- Guardrail trigger rate
- Evaluation sample failures
- Retrieval freshness
- Queue depth and escalation rate
- Token usage anomalies
Alerts tie to runbook sections. On-call should not grep logs without a map.
Incident tiers
Tier 1 — ops: stuck queue, single tool degraded, use runbook, escalate if unresolved in N minutes.
Tier 2 — engineering: model regression, deploy fault, data pipeline break, needs code or config change.
Tier 3 — vendor/provider: upstream model outage, cloud region issue, coordinated comms.
Roles documented in the runbook and in the contract for managed run.
Managed run versus client on-call
Some clients want Incubics on the pager for evaluation failures, cost anomalies and first-response incident handling — monthly fee, systems under management.
Others take full ops after handover. We deliver runbooks, dashboards, training and optional hypercare window.
Both models work. The failure mode is unstated assumption.
Air-gap, OT and special cases
Manufacturing and energy sites may have no overnight IT on site. Edge deployments and delayed sync change what "incident" means.
Runbooks account for who can act locally, what can run offline and how updates propagate when connectivity returns.
The human cost of missing ops
Incidents without runbooks burn senior engineers on repetitive triage. Users lose trust faster from slow messy recovery than from a brief controlled outage with comms.
FinOps incidents — runaway spend — need owners too, not only functional outages.
Dress rehearsal checklist
Before production traffic:
- Simulate tool outage — does escalation work?
- Simulate bad deploy — rollback tested?
- Page on-call — did they find the dashboard?
- Walk finance through cost alert
One afternoon prevents a week of reputation repair.
Who is awake at 3am — the steering committee question
Ask it in the go-live readiness review. Accept only named answers: person, team, contract clause, backup.
If Incubics managed run is the answer, it is in the statement of work with scope and response times. If client IT is the answer, handover proves they can execute the runbook.
Hypercare after go-live
First two to four weeks warrant elevated support: daily check-ins, fast harness fixes, tuned alerts. Hypercare is scoped in the increment — not infinite, not zero.
Exit hypercare when metrics stabilise and on-call executes runbook without vendor rescue.
Post-incident review
Every severity-one incident gets blameless review: timeline, root cause, harness gap, runbook gap, action items with owners. Findings feed evaluation and documentation.
Skipping post-incident review guarantees repeat failures.
Dependencies and third parties
Runbooks name upstream dependencies — model provider, identity, database, retrieval index — and expected communication channels when they fail. Agents fail cascaded more often than in isolation.
Cost incidents are incidents too
Runaway inference spend triggers the same discipline as functional outage: detect, contain, diagnose, fix, post-incident review. FinOps alerts belong on the same ops calendar as error-rate alerts.
Containment may mean disabling an expensive tool route, throttling concurrency or falling back to draft-only mode — pre-agreed in runbook, not invented under stress.
Contracting managed run clearly
If Incubics holds the pager, response times, scope — evaluation fixes versus feature work — and escalation to client IT are in the statement of work. Ambiguous SLAs cause friction at the worst time.
If client IT holds the pager, hypercare and paired drills before exit prove they can execute rollback without vendor rescue.
Part of Deliver and Run
Deliver includes runbook handover and ops training — not only developer docs.
Run is the ongoing option: evaluation, upgrades, supervision, incident first response, quarterly roadmap.
The method names Run as the fourth movement for a reason. Production without run is a handoff to nowhere.
Building operational maturity before scale
Operational maturity scales with traffic. Launching to ten users hides weak runbooks; launching to ten thousand exposes every gap. Increment one production release should include operational readiness equal to functional readiness — not "we will harden later when volume grows."
Capacity planning for inference and tool backends belongs in pre-go-live checklist. Load test at expected week-four volume, not demo volume.
Escalation trees should name backups. Primary on-call on vacation should not mean vendor becomes implicit pager forever.
Run cost anomalies often precede functional incidents — looping agents, cache misconfiguration. Train on-call to treat FinOps alerts as first-class.
Joint ops-engineering weekly review in the first month catches runbook gaps while context is fresh.
Documentation lives in the same repo as runbooks where possible — versioned, reviewed, not a SharePoint orphan.
When clients take ops in-house, shadow week is minimum — client leads incident with vendor coaching.
Managed run SLAs should distinguish severity: user-facing workflow stopped versus degraded quality versus internal-only staging issue.
Quarterly disaster rehearsal — rollback, provider failover, manual queue mode — keeps skills sharp.
Operations is not the appendix of AI. It is the proof the system is real.
Read How we work, Engagements and FAQ for process and commercials. Select your region — India, Middle East, ANZ or Other — when you write.
Closing note
Steering committees should ask the 3am question in writing before sign-off on go-live. Accept only named owners and tested runbooks as answers.
Production AI is a programme of small disciplined choices: discovery before build, evaluation before users, APIs before agents, runbooks before scale, adoption alongside code. Skip any one and the demo survives while the business outcome does not.
Incubics works that way by design — two-week discovery, production by week twelve, optional managed run. Offices in Bengaluru, Pune and the USA. Legal entity IQLEXA Technologies Private Limited.
If this piece matches a problem you are living, the next step is a conversation, not another pilot.
Write the 3am answer down before go-live. If it is not written, it is not ready. Test the pager before users depend on the system.
hello@incubics.com if go-live is approaching and nobody has answered the 3am question yet. Discovery and build can still close the gap before users find it for you.