Home › Insights › Blogs › Data & AI

How to Connect a Language Model to Your Own Data Without Leaking It

Prashant Talesara Prashant Talesara
Last updated: 9 Oct 2026
Get an AI summary of this post on Perplexity ChatGPT Gemini

You connect a language model to private data by retrieving from it, not training on it. At query time the system fetches only the records the asking user is already permitted to see, passes those to the model as context, and keeps the source data inside your existing access controls.

There is a conversation happening in most enterprises right now where both sides are correct and nobody can move.

The business wants a language model that can answer questions about the actual company. Legal and risk will not release the data those questions depend on. Neither position is unreasonable. A model that cannot see your contracts cannot answer questions about your contracts, and a risk function that waves through an unreviewed data flow into a third-party system is not doing its job.

The deadlock usually breaks on a false assumption: that giving a model access to data means giving the data to the model. It does not, and the difference is architectural. The design that makes the answers useful is the same design that keeps the data contained, which means this is a reviewable engineering decision rather than a question of how much risk appetite anyone has.

A general-purpose model knows a great deal about the world and nothing whatsoever about your business. It has never seen your master services agreements, your claims history, your patient records, your support tickets, or the three internal wikis where the real operating knowledge lives.

Which means every question actually worth asking is a question it cannot answer. Not “what is a force majeure clause” but “which of our active contracts have one, and what do they exclude.” Not “how do I process a claim” but “why was this claim rejected and what does our policy say about appeals.” The value lives entirely in the proprietary half.

That proprietary half is precisely what sits behind access controls your organisation spent years building. Role-based permissions, record-level restrictions, a decade of carefully granted and revoked access. It is governed data, and it is governed for good reasons.

So the pilot that impressed the steering committee ran against a sanitised extract, and the production version needs the real thing. That is the gap where most enterprise AI programmes stall, and it is not a technical failure. It is a design question that was deferred until it became a blocker.

Retrieval with permission filtering, explained

Retrieval-augmented generation is a design in which the model answers from documents fetched at the moment of the question, rather than from anything it absorbed during training. The data is supplied as context for a single response and is never written into the model. Add permission filtering to that design and the model can only ever be given content the person asking is already cleared to read.

The sequence matters, so here it is in order.

  1. Indexing. Your documents are broken into passages and indexed for search. Critically, each passage carries forward the access permissions of the system it came from, as metadata attached to the passage.
  2. The question arrives. A named, authenticated user asks something.
  3. Entitlement resolution. The system establishes who that person is and what they are permitted to see, from your existing identity provider rather than a separate list maintained by the AI system.
  4. Filtered retrieval. The search runs against only the passages that user can access. Everything else is not ranked lower. It is absent.
  5. Context assembly. The handful of permitted passages that best match the question are placed into the prompt.
  6. Generation. The model composes an answer from that context.
  7. Citation. The response points back to the source documents, each of which the user can open in the originating system, because they had access to it all along.

The single most important detail is in step four. The filter runs before retrieval, not after.

This sounds like a small implementation choice and it is the entire control. If a system retrieves broadly and then filters the output, the model has already read restricted content. It may have summarised it, inferred from it, or carried its substance into an answer that contains no quotable fragment but plenty of information. Post-filtering protects the text and leaks the meaning. Pre-filtering means the restricted content never existed as far as that query is concerned.

When you review a proposed design, this is the question to ask first, and the answer should be specific. “Permissions are enforced” is not an answer. “The retrieval query is constrained by the user’s entitlements before the search executes, server-side” is.

Where data crosses a boundary and where it does not

Most of the anxiety in these reviews comes from not knowing where the data physically is at each step. Here is the map.

StepLeaves your control?What to verify
Source documents at restNoThey never move. The index references them.
Search index and embeddingsOnly if hosted externallyWhere it runs, and that it is treated as in-scope data
Entitlement checkNoRuns server-side, before retrieval, against your identity provider
Retrieved extract plus questionYes, if the model is third-partyRegion, retention period, no-training commitment
Generated answerSame boundary as aboveNothing new crosses. The answer returns the way it came.
Prompt and response logsDepends entirely on configurationWhat is captured, where it is stored, how long it is kept

Only one row is an actual crossing, and only when the model is a third-party service. Run the model inside your own cloud tenancy or on your own hardware and nothing leaves at all, at the cost of operating it yourself.

Three things on that table surprise people in review, and they are worth naming because they are where real exposure hides.

The embeddings are derived from your data. A vector is not plaintext, and it is also not anonymised. Published research has demonstrated that meaningful portions of source text can be reconstructed from embeddings. Treat the vector index as holding the same classification as the documents it was built from, not as a harmless numerical artefact.

The logs are usually the largest uncontrolled copy. Prompt and response logging is on by default in a lot of tooling, and it accumulates verbatim extracts of exactly the sensitive content you were careful about everywhere else. It is the most common way a well-designed system leaks, and it is a configuration problem rather than an architecture problem, which is why it gets missed.

Deletion has to reach three places. When a record must be removed, it has to go from the source system, from the index chunks derived from it, and from the embeddings. A deletion path that only clears the source leaves the content answerable through retrieval.

Why fine-tuning is the riskier answer

Someone in the room will ask why you do not simply fine-tune a model on the company’s data. It sounds tidier, and the answer is that it moves the data somewhere you cannot get it back from.

Fine-tuning writes information into the model’s weights. That has four consequences a risk function should weigh carefully.

  • It cannot be revoked per record. There is no delete. Honouring a deletion request against a fine-tuned model means retraining it without that record, which is a project rather than an operation.
  • It can be extracted. Models can reproduce training data, sometimes verbatim, under the right prompting. The research on memorisation and extraction is well established, and the risk rises with how distinctive the record is, which is exactly backwards from what you want for sensitive data.
  • Permission filtering becomes impossible. Weights have no concept of who is asking. A fine-tuned model is a single permission tier, which means it can only safely be trained on data that every one of its users is entitled to see. In most enterprises that set is nearly empty.
  • It is stale from the day it finishes. The model knows the corpus as of training. Retrieval reflects the document as it is now.

None of this makes fine-tuning illegitimate. It is genuinely good at teaching a model tone, output format, and domain vocabulary, and those are the things it should be used for. The error is reaching for it to teach a model facts, because facts belong in a retrieval layer where they stay current, stay revocable, and stay governed by the permissions they already carry.

Put plainly: retrieval keeps your data as data. Fine-tuning turns your data into model.

Residency, logging and what auditors ask for

Once the design is settled, the review turns to where things physically sit and what you can prove. In a retrieval system the data has three resting places, and a complete answer accounts for all three rather than just the first.

The source systems, which you already know and already document. The search index and its embeddings, which is new infrastructure and often the one nobody assigned a classification to. And the logs, which are new, high-volume, and contain the sensitive content in the clear.

Residency applies to all three independently. It is entirely possible to keep source documents in-region, host the index in-region, and then send extracts to a model endpoint in another jurisdiction, or write logs to a monitoring service whose storage region nobody checked. The question to put to the architecture is not “is this in region” but “name the region for each of the three, and for inference.”

When an audit comes, in our experience the requests are consistent and none of them are exotic:

  • A data flow diagram showing every point the data moves and rests, including the index and the logs.
  • The agreement with the model provider, specifically the no-training clause, the retention period, and the processing region.
  • Evidence that permission filtering is enforced server-side, rather than in a client that could be bypassed.
  • The logging configuration, with retention periods, and who can read the logs.
  • A deletion procedure that demonstrably reaches the source, the index and the embeddings.
  • Access records for the index itself, which is a new system that somebody can query directly.

A team that built this deliberately can produce all six in an afternoon. A team that bolted it together will struggle on the third and the sixth, which is a useful diagnostic in itself. Governing this well sits alongside the broader control framework covered in our AI governance framework for regulated industries, which deals with who decides and signs off rather than how the data moves.

What to require from any vendor proposing this

If someone is pitching you a language model connected to your internal data, these are the eight things to get in writing before approval. They are deliberately specific, because each one corresponds to a failure that happens in practice.

  1. Permission filtering enforced server-side, before retrieval. Not in the application layer, not after the fact.
  2. A written no-training commitment covering your prompts, your documents and the outputs.
  3. A stated retention period for prompts and responses, with zero retention available for the most sensitive categories.
  4. A named processing region for inference and for log storage. “Global” is not a region.
  5. Embeddings treated as in-scope data, classified and controlled like the source text.
  6. Logging controls you configure, so you decide what content is captured rather than inheriting a default.
  7. Per-record deletion that reaches the index, with a stated completion time.
  8. Evidence an auditor will accept, meaning the diagram, the agreement and the access records, not a confidence that it is handled.

A vendor who has built this before will answer all eight without flinching, because they have been asked before. One who becomes vague around three, five and seven is describing a demo rather than a system. That is a useful filter to apply early, and it is the same filter we apply to our own generative AI development work before anything touches regulated data.

Also read

This article covers the architecture. For the question of what this costs to build and operate, that is covered separately in the cost breakdown, and the wider practice sits under our AI solutions.

The bottom line

The disagreement between the business and the risk function is usually built on a premise neither side has examined: that useful means exposed. It does not. Retrieval with permission filtering lets a model answer from data it is never given, under the access controls you already run, with citations that make every claim checkable.

What that leaves is a short, concrete review. One boundary crossing to govern. Three resting places to assign a region and a classification. One filter to confirm runs server-side and early. Eight commitments to get in writing.

That is a conversation a compliance lead can finish, which is more than can be said for the one most organisations are currently having.

Get your AI data architecture reviewed

Bring us the design on the table and the constraints you are working under. We will map where the data actually crosses, what your auditors will ask for, and what to fix before approval.

Book a Free Call
#RAG #Generative AI #Data Privacy #AI Compliance #AI Governance #Enterprise AI
Share

Frequently asked questions

How do you connect an LLM to private company data?
Through retrieval, not training. Your documents are indexed, and each piece keeps the access permissions of the system it came from. When someone asks a question, the system resolves who they are, filters the index down to only what that person is already allowed to see, and passes those extracts to the model as context for that one answer. The model reads the context, writes a response, and retains nothing. The source data stays in your systems under the access controls you already operate, which is what makes the approach reviewable.
Does using a language model mean my data trains the model?
Not with retrieval, and not with any reputable enterprise provider by default. In a retrieval design the data is supplied as context at query time and is never written into the model's weights, so there is nothing to untrain. What varies between providers is retention of the prompts and responses themselves, which is a contractual matter rather than an architectural one. Get a written no-training commitment and a stated retention period, and verify both in the agreement rather than the marketing page.
What is retrieval-augmented generation (RAG)?
Retrieval-augmented generation is a design in which a language model answers using documents fetched at query time rather than only what it learned during training. The system searches your indexed content for passages relevant to the question, inserts them into the prompt, and the model composes an answer grounded in those passages. Two consequences matter for risk. Answers can cite their sources, so claims are checkable. And because the data is supplied per query rather than absorbed into the model, it stays governed by your existing access controls.
How do you stop an LLM showing a user data they should not see?
Filter before retrieval, not after. Each indexed passage carries the permissions of its source system. When a question arrives, the system resolves the user's entitlements and restricts the search to passages that user can already access, so unauthorised content is never retrieved and never reaches the model. The weak pattern is retrieving broadly and filtering the answer afterwards, because by then the model has already read the restricted content and may have summarised it. Insist the filter is enforced server-side, before the search runs.
Where does my data actually go when a model answers a question?
In a well-built retrieval system there are three resting places and one crossing. The source documents stay in your systems. The search index and its embeddings sit wherever you host them. Prompt and response logs sit wherever logging is configured. The crossing is the retrieved extract plus the question, which travels to wherever the model runs. If the model is self-hosted or runs inside your cloud tenancy, nothing leaves your boundary at all. If it is a third-party API, that extract is the only thing that crosses, and its handling is governed by your agreement.
What should I require from a vendor proposing an LLM on our data?
Eight things, in writing. Permission filtering enforced server-side before retrieval. A no-training commitment covering your data. A stated retention period for prompts and responses. A named processing region. Confirmation that embeddings are treated as in-scope data. Logging controls you configure rather than inherit. Per-record deletion that propagates to the search index. And evidence you can hand an auditor, meaning a data flow diagram and access control records rather than assurances. A vendor who cannot answer all eight has not built this before.
Prashant Talesara
CTO, Kansoft

CTO at Kansoft. 18 years of experience building data, AI, and agentic systems for global enterprises across healthcare, financial services, and industrial sectors.

Related articles

Need help with your next project?

Our engineering experts can help you build something exceptional.

Book a Free Call