AI ON BUSINESS DATA
What the LLM sees
Not your database, not your index, not your collections. The question, the names it needs to map it, and the result it writes up. That is the list.
The question behind the question
"What does the model see?" is the first thing a security review asks about an assistant over business data, and it deserves a precise answer rather than a reassuring one. The precise answer is a short list, and this page is that list. It describes the shape of the design - what has to cross to the model for the assistant to work, and what never needs to - and it is the shape to hold any vendor to.
What it sees
The question, as typed. There is no way to understand a question without reading it.
The vocabulary. The names of the collections - orders, products, customers, invoices - and of their fields, with a line describing each: what region means, that status takes one of five values, that order date is a date. This is metadata, not data. It is what lets the model turn "orders by region this month" into a request against the orders collection, faceted by region, filtered to the month, without guessing field names.
The result it is writing up. When the index has answered - North 395 · West 247 · East 168 · South 142 - the rows and counts that make up that answer are given to the model so it can write the summary: North leads order volume by a wide margin. The result is the same thing the user is about to see on screen. The model sees the answer to this question, not the records the answer was counted from.
The conversation so far, so that "now only North" can be understood as a narrowing of the previous question.
- The question, as typed
- The names and descriptions of collections and fields, so it can map the question
- The result it is asked to write up: the rows and counts that answer this question
- The conversation so far, for follow-ups
- The production database
- The index, or any collection as a whole
- Records the question did not retrieve
- Access keys, credentials, the schedule
- Anything from a question another user asked
What it never sees
The production database, because the assistant never reads it: it reads an indexed copy. The index, because the model never searches; it asks, and the index searches. Any collection as a whole, because no request returns one. Records the question did not retrieve, because only the result of the request is written up. Credentials, keys, refresh schedules, access lists, because none of them is part of any request. And nothing from a conversation another user had, because each conversation carries only its own context.
| Item | Sent to the model? | Why |
|---|---|---|
| The question | Yes | It has to understand what is being asked |
| Collection and field names, with descriptions | YesThe vocabulary, not the data | It maps the question onto fields that exist |
| The result for this question | YesRows and counts, as returned | It writes the summary from what came back |
| The conversation so far | Yes | Follow-ups depend on it |
| The whole collection | No | The index does the searching; the model never scans |
| The production database | No | It is never on the path |
| Records outside the result | No | Only what the request retrieved is written up |
| Keys, schedules, access lists | No | Not part of any request |
Why the result has to cross
A reasonable question is whether the model could be kept away from the numbers altogether, receiving only the shape of the answer and having the figures filled in afterwards. It could, for a template. It cannot, for a summary: a sentence like North far ahead with 395 orders, followed by West with 247 is written by reading the result, and a model that has not read the result cannot write it. So the honest statement is the one on this page: the retrieved result for the question crosses, the dataset does not. The design choice that keeps that bounded is the size of the result - a facet with four regions is four numbers, not 952 orders - and the fact that the request, not the model, decides what is retrieved.
Under whose account
What crosses, crosses to a provider you chose, under your own account and your own key, and is handled under your agreement with that provider, not under a vendor's. That matters because the provider's terms are the terms that govern what it may retain or train on, and because an account you hold is an account you can cap, audit and close. Bring your own LLM covers the arrangement; Data flow map places this flow beside the three others.
How to verify it
Ask the vendor for the list on this page, in this form: what is sent, what is not, and under whose account. Ask what bounds the size of a result. Ask whether the model can issue a query of its own or only choose from a vocabulary, because a model that can write SQL can retrieve whatever the SQL retrieves. And ask whether the same questions have the same answers for every product in the platform, because they usually do not, and the honest vendor will say so.