Capabilities
Retrieval and tool integration
What it is
Retrieval spans vector search, keyword search, graph traversal and live queries against warehouses — unified behind policies that respect tenant, role and data residency. Tool integration wraps your APIs and databases as typed functions agents can call, with validation, retries and circuit breakers.
This capability is rarely visible to end users but determines whether production systems survive real load and real permissions.
Design decisions
- Chunking and metadata strategy per content type.
- When to embed versus when to query structured data directly.
- Reranking and citation assembly for user trust.
- Tool schema design — small, composable functions beat mega-endpoints.
- Auth propagation — user context versus service accounts.
- Caching and freshness SLAs for expensive queries.
When it pays back
Every assistant and agent investment depends on this layer. Payback shows up as higher answer accuracy, fewer tool errors and lower inference cost when retrieval replaces long context windows.
It is also the enabler for incremental rollout: add a new tool, add evals, promote to production without rebuilding the whole agent.
Common triggers
- Pilots worked on static exports but fail on live data.
- Agents timeout calling slow legacy APIs.
- Security blocked production because service accounts were over-permissioned.
- Duplicate retrieval stacks sprawl across teams.
How Incubics engineers it
Discovery: two weeks mapping systems, access patterns and latency. Week twelve: production retrieval and tool layer for the first use case, reusable for the next.
Perceive — weeks 1–2
We catalogue sources, API maturity, rate limits and identity models. We prototype retrieval quality on sample questions and tool reliability on non-production sandboxes. Output: integration architecture, security model, eval plan and fixed scope.
Engineer — weeks 3–10
We build ingestion and index pipelines, hybrid retrieval services and tool adapters with OpenAPI or internal SDK patterns. Observability traces each retrieval and tool call. Load tests match expected peak concurrency.
Deliver — by week 12
Production endpoints wired to the first agent or assistant, documented schemas, secret rotation procedures and runbook for index rebuilds and API degradation modes.
Run — ongoing
We monitor retrieval precision proxies, tool error budgets and cache hit rates. New sources and tools onboard through the same eval and security checklist.
Failure modes
- Embedding stale PDFs while live pricing sits in SQL nobody connected.
- Tool schemas that accept free-text where enums are required.
- Retrieval returning cross-customer chunks — tenancy bug, not model bug.
- No timeout strategy — one slow ERP call freezes conversations.
- Logging full payloads including secrets.
Teams sometimes over-index on vector databases when a simple indexed query would be faster and exact. We choose mechanisms per question type.
What you get
- Retrieval service with hybrid search and reranking.
- Ingestion pipelines and index lifecycle automation.
- Tool gateway with auth, validation, rate limits and audit.
- Developer documentation and contract tests for each tool.
- Tracing integrated with your APM or OpenTelemetry stack.
- Performance benchmarks and capacity guidance.
- Patterns for adding sources without fork-lifting the agent.
What we refuse to ship
We refuse tools with blanket admin credentials. We refuse retrieval without tenant filters when data is multi-customer. We refuse undocument tool behaviour that agents depend on.
Do we need a vector database?
Not always. Discovery matches store choice to query patterns. Many systems combine warehouse SQL, search engines and vectors.
How do you handle API rate limits?
Queuing, backoff, caching and pre-fetch where business rules allow. Agents receive graceful degradation messages, not infinite retries.
Can tools call mainframe systems?
Through stable API layers — see legacy integration. Direct terminal scraping is not a production strategy we endorse.
How is retrieval tested?
Question-answer pairs with expected source IDs, plus manual review samples. Metrics track recall of correct chunks on held-out sets.
Is this reusable across multiple agents?
That is the design intent. Shared tool and retrieval services reduce duplication and centralise security review.
What about real-time data?
We define freshness requirements per field. Some answers pull live; others use indexed snapshots with visible timestamps in the UI.
Next step
Start with two weeks.
A fixed-fee discovery gives you a ranked use-case portfolio, a target architecture, a cost model and a build proposal you can take to your board. If we don't find a case worth building, we tell you.