Built from parts you can name
We work with mainstream model providers, ordinary vector and keyword search, and your own systems through their APIs. Nothing here is exotic, and that is deliberate: the parts have to be supported by your team after we leave.
How a request is handled
Access is applied before retrieval, so the assistant can never return a passage the person asking is not allowed to open. The source check is a separate step from generation: if the citations do not support the answer, it is not shown.

Where it runs
Deployments are containerised and described as code, so the same build runs in your cloud tenant, on your own hardware or in ours. Test and production are separate environments, and the evaluation suite runs against both before anything is promoted.
Where a model has to run inside your perimeter, it is deployed the same way — with the accuracy comparison from the pilot attached to the decision.
The parts
Models
Models from the main providers for hosted work, open-weight models inside your own environment where data may not leave it. The choice follows the constraint.
Retrieval
Vector and keyword search over your documents, hybrid ranking, permission-aware filtering, citation back to the source page.
Orchestration
Explicit pipelines and tool-calling agents with step logging, retries, timeouts and a defined hand-over point to a person.
Evaluation
Labelled question sets from your own operation, retrieval and answer metrics tracked per release, failure cases kept as tests.
Integration
APIs and web services against the systems you already run; no rip-and-replace, and no second copy of your data where it can be avoided.
Delivery
Containerised services in your cloud account or ours, infrastructure as code, separate test and production environments.
Observability
Per-request logs of prompt, retrieved passages, model version and outcome, retained according to your policy rather than ours.
Access control
Your identity provider, your groups and roles; the assistant never has wider access than the person asking.
Where the data lives
- In your cloud tenant, on your hardware, or in ours — decided before any code is written.
- Documents are indexed where they already are, with no second copy unless the architecture requires one.
- External model calls only for the material your policy allows; everything else stays inside the perimeter.
What is logged
- Per request: the question, the passages retrieved, the model version, the answer and the outcome.
- Retention set by your policy, not ours; the log is exportable for an audit.
- No training on your data, and no use of it to improve a vendor model.
When data may not leave
Open-weight models run inside your perimeter, with the accuracy trade-off stated in writing before the work starts. This is a real cost, not a checkbox — the pilot measures both options on your own cases so the decision is made on numbers.
What you get at hand-over
- Infrastructure as code and a deployment procedure your team can run.
- The evaluation suite, so quality can be re-measured after any change.
- Runbooks, monitoring and the data-flow description used for review.