AI Vendor Evidence Gap Notes #1
A no-training claim may be true.
It may be useful.
It is still not evidence-complete.
“We do not train on your data” is one of the most common AI vendor statements. It sounds as though it answers the buyer’s main concern.
But it usually answers only one part of the review question.
The real question is not only:
Will the vendor train a model on our data?
The better question is:
After our data enters the AI product, what happens to it?
That means looking past the no-training statement and asking about the wider data path:
prompts and outputs;
uploaded files;
logs and metadata;
retention;
human and support access;
subprocessors and model providers;
deletion;
audit logs;
contractual scope.
A vendor may not train on customer data and still retain prompts, generate operational logs, allow support access, collect metadata, route content through another provider, or preserve records for safety and abuse monitoring.
That is the evidence gap.
Claim
The vendor says:
“We do not train on your data.”
Or:
“Customer data is not used to train our models.”
Or:
“Your prompts and outputs are not used for model training.”
These statements sound similar.
They may not mean the same thing.
One may cover broadly defined customer data. Another may cover prompts and outputs only. Another may address model training while saying nothing about retention, logging, support access, safety review, evaluation, telemetry, or product improvement.
The wording matters.
So does the source.
A public FAQ is not the same as a signed contract.
A product page is not the same as a data processing addendum.
A company-level statement is not automatically evidence for a particular product, plan, workspace, endpoint, region, configuration, or customer agreement.
Why it sounds sufficient
The claim answers the question many buyers have been trained to ask:
Will the vendor use our data to train its model?
That is a legitimate concern.
If proprietary, customer-confidential, regulated, or sensitive business data may be used for model training, the buyer needs to understand that before allowing real data into the product.
When the vendor says no, the main issue can appear closed.
That is where the review can become too narrow.
A no-training claim addresses one possible use of customer data.
It does not explain the operational data path.
Training, retention, logging, safety monitoring, support access, human review, model routing, and deletion are separate questions.
What it actually proves
A current, product-specific no-training statement may support a narrow conclusion:
The vendor states that defined customer data is not used to train or improve its models.
That is meaningful evidence.
Its strength still depends on scope.
The buyer should be able to determine:
which data types are covered;
which product, plan, and feature are covered;
whether the commitment applies by default or only after configuration;
whether customers can opt in;
where the commitment is written;
whether the source is public, contractual, audited, or customer-specific;
whether the statement applies to the buyer’s workspace, region, endpoint, and agreement;
when the evidence was last checked.
Without those details, the claim may be reassuring but too broad to support the intended workflow.
A public evidence example
The CodeYourCompliance public evidence profile for OpenAI records several public evidence surfaces, including OpenAI’s business-data documentation, Enterprise Privacy page, Trust Portal, contractual materials, and product-specific data-control references.
The profile is an evidence index.
It is not verification of the settings, enabled features, contract terms, or actual data path for a particular customer workspace.
OpenAI’s business data privacy page states that, by default, inputs and outputs from specified business products—including ChatGPT Enterprise, ChatGPT Business, ChatGPT Edu, ChatGPT for Healthcare, ChatGPT for Teachers, and the API platform—are not used to train or improve OpenAI’s models.
That is a meaningful public commitment.
It establishes a training-use boundary for defined business offerings by default.
It does not describe the full lifecycle of the data.
OpenAI’s separate Enterprise Privacy page describes additional service-specific boundaries.
For example, it distinguishes retention controls across products. It states that business data may be subject to human review on a service-by-service basis. It also explains that, except for certain endpoints and features, API inputs and outputs may be retained for up to 30 days for service delivery and abuse detection, while eligible customers may request zero data retention for qualifying API use cases.
These statements are not necessarily contradictory.
They answer different questions.
The no-training statement addresses whether defined business data is used to improve the vendor’s models.
The retention statement addresses how long certain operational data may remain.
The human-review statement addresses who may access stored data and under which service-specific conditions.
The API documentation addresses endpoint- and configuration-specific handling.
OpenAI’s data-use guidance also distinguishes business products from services for individuals. Business-product inputs and outputs are excluded from training by default, while individual users may have different defaults and training controls.
That distinction matters.
A company-level slogan should not be applied automatically across business and personal services, different products, plans, workspaces, endpoints, features, or configurations.
The public evidence supports a defined no-training conclusion.
It does not, by itself, establish the complete operational data path for the buyer’s intended workflow.
What it does not prove
A no-training claim does not automatically establish:
Retention. Prompts, outputs, files, or other records may still be retained for service delivery, safety, abuse monitoring, support, or legal requirements.
Logging and metadata. Operational logs, account records, security events, classifier results, and usage metadata may follow different rules from customer content.
Human access. Employees, contractors, support personnel, safety reviewers, or other authorized roles may be able to access certain stored data under defined conditions.
Model-provider routing. The claim may not show whether another model provider or infrastructure service receives prompts, retrieved context, files, outputs, or metadata.
Subprocessor handling. The vendor may use third parties for infrastructure, monitoring, support, analytics, transcription, model inference, or abuse review.
Storage-layer coverage. Application history, caches, indexes, embeddings, backups, support copies, and safety records may not share the same lifecycle.
Deletion. Deleting visible conversation history does not automatically establish deletion from logs, backups, secondary systems, or downstream providers.
Auditability. The claim does not show which events are logged or whether the buyer can reconstruct access, routing, deletion, configuration changes, or human review.
Contract scope. A public page does not necessarily provide the same commitment as a DPA, service agreement, order form, product addendum, or negotiated term.
None of these gaps proves that the vendor is using customer data improperly.
They show that the training claim is narrower than the decision the buyer is trying to make.
Weak-answer pattern
This is the narrow-promise pattern.
The buyer asks how customer data is handled.
The vendor answers:
“We do not train on customer data.”
The answer may be true.
But it addresses training while leaving the operational data path unclear.
The review becomes weak when the file records:
No customer data used for training.
and then treats the wider data-handling question as resolved.
Training use has been answered.
Retention, logging, support access, human review, model routing, subprocessors, deletion, and contract scope have not.
Evidence request
Do not ask the vendor to repeat the slogan.
Ask for the data path.
A stronger request is:
Please provide product- and plan-specific evidence showing how prompts, outputs, uploaded files, logs, metadata, diagnostic records, safety events, support access, subprocessors, and model providers are handled. For each data type, identify its purpose, storage location, retention period, deletion process, human-access boundary, provider path, customer control, and supporting contractual source.
Then ask:
Which data types are covered by the no-training commitment?
Does the commitment apply by default, or does it depend on settings or opt-in status?
Are prompts, outputs, files, metadata, logs, and support records treated differently?
Can employees, contractors, or third parties access stored content?
Which model providers and subprocessors receive customer data?
What retention and deletion rules apply to each storage layer?
Which product, plan, endpoint, region, and contract terms support the answer?
That is how a no-training claim becomes reviewable.
Review note
A weak review note says:
The vendor does not train on customer data.
A stronger review note says:
The vendor states that defined customer data is not used for model training. This supports a narrow training-use conclusion, but the available evidence does not establish the complete operational data path for the intended product, plan, workspace, endpoint, feature, and configuration. Retention, logging, metadata, human review, support access, model-provider routing, subprocessors, deletion, auditability, and contractual scope require separate evidence.
That note does not accuse the vendor of making a false statement.
It prevents a narrow statement from being treated as a complete review record.
Usage boundary
Until the remaining data path is evidenced:
Limit use to low-sensitivity internal or test data. Do not introduce customer-confidential data, regulated data, source code, employee records, privileged business material, or externally relied-upon outputs where uncertainty about retention, human access, logging, provider routing, deletion, or contractual scope would materially change the appropriate usage boundary.
That is not approval.
That is review preparation.
The review unit is:
vendor + product + plan + use case + data type + region + contract terms + evidence date.
A no-training commitment can support that review.
It cannot replace the complete data path.
Bottom line
“We do not train on your data” may be true.
It may be useful.
It is still not the review.
The buyer still needs to know:
what enters the product;
where it goes;
how long it remains;
who can access it;
which third parties receive it;
what can be deleted;
what can be audited;
which commitments apply to the actual workflow.
Vendor claim is not evidence-complete until the claim is mapped to its source, scope, exceptions, and operational boundary.
Boundary
This material is for evidence structuring and review preparation. It does not provide legal, regulatory, audit, procurement, certification, or implementation advice.
The goal is not to approve or reject a vendor.
The goal is to make the evidence gap visible before real data use.
Reply if you want the sanitized sample evidence gap memo.
#AIVendorRiskAssessment #AIDataUse #AIGovernance #ThirdPartyRisk #DataPrivacy


