Capabilities
Model selection and evaluation
What it is
Different steps in a system need different models: a small fast classifier for routing, a larger model for extraction, an embedding model matched to your domain language. Selection runs structured benchmarks — offline on labelled sets, online in shadow — measuring quality, latency, throughput and unit cost.
Evaluation continues after selection. Every prompt change, fine-tune or vendor upgrade reruns the same suites before production promotion.
What we measure
- Task accuracy against business-defined rubrics, not generic benchmarks.
- Tool-call correctness and format adherence.
- Refusal and safety behaviour on your red-team set.
- Latency at p50 and p95 under expected concurrency.
- Cost per successful transaction at expected token volumes.
- Operational fit: residency, private deployment, support for required modalities.
When it pays back
Payback is immediate when teams debate models for months while pilots stall. A two-week benchmark on real tasks often collapses the option space to one or two viable stacks.
It also prevents expensive overruns — flagship models on every call when a smaller model handles most traffic — and prevents compliance surprises when a chosen vendor cannot run in your region.
You need this when
- Leadership asks which model without data.
- A vendor renewal is approaching and alternatives exist.
- Open-weight deployment became a security requirement mid-project.
- Quality regressed after a silent provider update.
How Incubics engineers it
Discovery includes model baselines in two weeks. By week twelve the production system runs the selected stack with automated eval gates — selection is not a one-off report.
Perceive — weeks 1–2
We define tasks, constraints and candidate models — frontier APIs, open-weight self-hosted, fine-tuned variants. We assemble eval datasets from production samples with redaction. Output: benchmark plan, acceptance thresholds, deployment constraints.
Engineer — weeks 3–10
We run benchmarks, analyse error modes and prototype routing — model cascades where cheaper models handle easy cases. Fine-tuning or distillation enters scope only if evals show clear lift worth the maintenance.
Deliver — by week 12
Production configuration documented with version pins, fallback models and rollback steps. CI blocks releases that fail thresholds. Stakeholders receive a readable scorecard, not raw logs.
Run — ongoing
We monitor quality and cost in production, re-benchmark when new models release and advise on upgrades quarterly. Your team owns the eval datasets; we maintain the harness.
Failure modes
- Choosing from public leaderboards unrelated to your tasks.
- Single-metric optimisation — cheap but unusably wrong.
- Ignoring inference architecture — model fits but cannot scale.
- Fine-tuning on dirty data — memorises errors.
- No rollback when a new version degrades tool calling.
Model religion wastes time. We document trade-offs and let evidence pick; we revisit when constraints or models change.
What you get
- Task definitions and labelled evaluation datasets.
- Benchmark reports with error analysis and recommendations.
- Routing configuration and version pinning strategy.
- CI-integrated regression suites.
- Cost projections at expected volumes.
- Runbook for model swap and emergency downgrade.
- Plain-language summary for non-technical approvers.
What we refuse to ship
We refuse a single-model edict without evals on your tasks. We refuse fine-tuning projects without baseline comparisons to prompt engineering and retrieval. We refuse hiding provider or version changes from ops.
Frontier or open-weight?
Whichever meets accuracy, latency, residency and cost thresholds together. Many systems use both — open-weight inside the VPC for sensitive steps, frontier for low-risk language tasks if policy allows.
How long do benchmarks take?
Initial comparative runs often complete inside discovery or early engineering. Full suites grow with the product but rerun in minutes to hours, not weeks.
Do you train custom models?
When evals prove fine-tuning or training beats alternatives on total cost of ownership. Otherwise we avoid unnecessary training debt.
Can we bring our own model contracts?
Yes. We integrate with your existing provider agreements and technical endpoints.
How do you handle multimodal needs?
Images, audio and documents each get task-specific metrics. Selection is per modality, not one model for everything.
Next step
Start with two weeks.
A fixed-fee discovery gives you a ranked use-case portfolio, a target architecture, a cost model and a build proposal you can take to your board. If we don't find a case worth building, we tell you.