How we work
Run
Shipping to production is not the finish line. It is the start of operational life. Language models change behaviour when vendors update weights. Your documents and processes change every quarter. Inference spend grows when usage grows unless someone watches it. Run is our managed operations service for AI systems we built or systems we inherit with a proper handover review.
What Run covers
Run is not generic IT helpdesk. It is AI-specific operations: keeping evaluation scores within agreed thresholds, investigating regressions, supervising agent actions against policy, managing prompt and model versions, controlling inference cost, and planning upgrades on a quarterly roadmap with your sponsor.
- Continuous evaluation against held-out and production-sampled data.
- Drift detection on inputs, outputs and tool success rates.
- Cost monitoring and FinOps recommendations — routing, caching, model swaps.
- Agent supervision: reviewing flagged sessions, tuning guardrails.
- Model and dependency upgrades with regression testing before promotion.
- Incident response within agreed SLAs.
- Quarterly roadmap session: what to improve, retire or expand next.
Evaluation in production
Pre-launch evaluation proves the system worked on known scenarios. Post-launch evaluation proves it still works on reality. We sample production traffic where policy allows, maintain golden sets aligned to business rules, and run scheduled regression suites on every prompt or model change. When scores drop, we diagnose: data drift, model drift, integration failure, or genuine new edge cases requiring product change.
Human review queues remain for subjective tasks — clinical notes, credit exceptions, complex tickets. Run staff tune sampling rates, coach reviewers, and feed corrections back into evaluation sets so the same mistake is caught automatically next time.
Cost control
Inference bills surprise executives when nobody owns them. Run includes monthly cost review: spend by use case, by model, by user cohort. We recommend concrete changes — smaller models for classification steps, batch processing overnight, aggressive caching of retrieval results — and implement approved optimisations. FinOps is part of operations, not a one-time architecture slide.
Agent supervision
Autonomous agents require oversight proportional to their blast radius. Run configures alerting on anomalous tool use, high-volume loops, or policy violations. We review sampled action logs with your compliance or operations team where required. Escalation paths defined in Deliver stay live and tested.
Model upgrades
Vendors release new models; open-weight checkpoints improve; your security team may mandate a move. Run manages upgrade as a controlled release: test in staging, run full regression, compare cost and latency, deploy with rollback ready. We do not auto-upgrade production on announcement day.
Taking it in-house
Run is optional. Every artifact needed to operate internally was delivered at handover. We support transition: paired operations weeks, shadowing, then reverse shadowing. Some clients keep Run for tier-two depth while internal staff handle tier-one user questions. No penalty for leaving.
Teams in Bengaluru, Pune and the USA provide follow-the-sun coverage where contracts require it. Coverage hours and SLAs are defined commercially, not assumed.
Reporting you can use
Run delivers regular reports executives and risk teams actually read — not raw log dumps. Monthly summaries cover evaluation trend, incident count and severity, cost versus budget, model and prompt changes promoted, and open remediation items. Quarterly roadmap sessions turn those trends into prioritised backlog for the next increment or internal team.
When an incident occurs, post-incident review documents root cause, blast radius, fix applied, and evaluation or guardrail changes to prevent recurrence. That record supports your operational risk process without us claiming to speak for your auditors.
When Run ends
Exit is planned. We agree success criteria for handover: internal team passes shadow on-call, regression suite runs green in your CI, cost dashboard owned by your FinOps. Run winds down over an agreed period — not an abrupt cliff on contract expiry.
Run complements your existing IT operations — we do not replace database administration or network ops unless scoped. We own the AI-specific layer: models, prompts, retrieval, agent behaviour, evaluation scores, inference spend.
Run versus internal MLOps hire
Hiring a full internal MLOps function takes quarters. Run covers the first operational year while you recruit — or permanently if you prefer opex over headcount. We document every runbook and pipeline so internal hire inherits working systems, not tribal knowledge.
Can you Run systems you did not build?
Yes, after a technical review. We need access to code, infrastructure, evaluation assets and documentation. Gaps become a short remediation scope before Run starts.
How is Run priced?
Monthly fee tied to systems under management, complexity and SLA tier. Detailed in the build proposal or a separate Run contract.
Does Run include new features?
Stabilisation and minor tuning are in scope. New use cases or major capability additions are fixed-scope increments, same as the original build method.
Next step
Start with two weeks.
A fixed-fee discovery gives you a ranked use-case portfolio, a target architecture, a cost model and a build proposal you can take to your board. If we don't find a case worth building, we tell you.