Skip to content
AI Metric

Chris M.Reviewed

Is It Safe to Upload Construction Documents to AI?

Take an illustrative case. A quantity surveyor has a 90 page tender to summarise before Monday. At half past four on Friday she pastes it into the chatbot she uses at home, gets a good summary in nine seconds, and goes to the pub. Nothing visibly happens. No alert fires, no policy is breached that anyone can name, and by Monday the only trace is a better summary than she would otherwise have written.

Whether it is safe to upload construction documents to AI tools depends almost entirely on the type of account used, and hardly at all on which company made the tool. On the consumer accounts of OpenAI and Google, content may be used to train models or be seen by human reviewers 1 13. On their commercial accounts, the same vendors state, and in one case contractually undertake, that it is not. That distinction carries more weight than any other fact on this page.

Three qualifications follow immediately, and they are where most business policies go wrong. No training is not the same as no retention. No retention is not the same as no human ever seeing it. And for a UK business, the reassurance you are most likely to be offered about where the data sits does not mean what it sounds like.

The line that actually matters

Take the vendors one at a time, in their own words.

OpenAI states that by default it does not train on any inputs or outputs from its products for business users, including ChatGPT Business, ChatGPT Enterprise and the API, and that when you use its services for individuals, such as ChatGPT, it may use your content to train its models 1. Anthropic states that by default it will not use inputs or outputs from its commercial products to train its models 4, and its Commercial Terms of Service go further and say that Anthropic may not train models on Customer Content from the Services 5. Microsoft states that prompts, responses and data accessed through Microsoft Graph are not used to train foundation models, including those used by Microsoft 365 Copilot 7. Google commits not to use customer data to train or fine-tune the models supporting its Workspace generative AI services without the customer's prior permission or instruction 12.

Now the consumer side of the same companies. Google's guidance to consumer Gemini users is unusually direct, and it is the most useful sentence available on this subject because it comes from the vendor rather than from a commentator: users are told not to enter confidential information that they would not want a reviewer to see or Google to use to improve its services 13.

If you take one thing from this guide, take that. The vendor is telling consumer users not to put confidential material in. A construction business that has not bought commercial accounts, and has not said so to its staff, is relying on people to have read a privacy page that tells them the opposite of what they assume.

One distinction is worth drawing out for anyone writing a policy. Some of these commitments sit in policy pages, which a vendor can revise, and some sit in contracts. Anthropic's no-training undertaking is in its Commercial Terms of Service 5, and Microsoft's position for organisational use is carried by its Data Protection Addendum and Product Terms 8. A contractual undertaking is a materially stronger thing to rely on than a published position, and asking procurement to establish which of the two a supplier is offering is a better question than asking whose marketing is most reassuring.

What "we do not train on your data" does not mean

Three gaps, each of which catches people out.

It does not mean nothing is kept. OpenAI's application programming interface (API) documentation states that by default abuse monitoring logs are generated for all API feature usage and retained for up to 30 days, unless longer retention is required by law or is reasonably necessary to protect the service or a third party 3. Google Workspace retention is administrator controlled 12. None of that is sinister and all of it is invisible to the person typing.

It does not mean no human sees it. For the Azure hosted models, Microsoft states that the models are stateless and that prompts and completions are not used to train or improve base models 11, and in the same document records that by default a sample of flagged prompts and completions is stored for human review by authorised Microsoft employees, with only customers approved for modified abuse monitoring avoiding that 11. On the consumer side Google states plainly that chats reviewed by human reviewers are not deleted when a user deletes their activity, and are retained for up to three years 13.

It does not undo what has already happened. Turning the training control off in a consumer account stops future use, and OpenAI is clear that conversations still appear in chat history when it is off 2. It is a forward-looking switch, not a recall. OpenAI documents one mode with a stated deletion period, Temporary Chat, which it says is deleted after 30 days and may still be reviewed to monitor for abuse 2. Somebody who used a personal account for a client's tender last March does not fix it by changing a setting today.

It does not survive a helpful click. Anthropic states that where a user explicitly reports feedback or bugs, for example via the thumbs up or thumbs down button, those chats may be used for training 4. That is an entirely reasonable arrangement, and it means a well-meaning staff member giving a thumbs up on a good answer can move a protected commercial conversation into a different category. If your policy does not mention the feedback buttons, it has a hole in it.

The residency question, and a correction

Ask a supplier where your data is processed and you will often be told it stays within the EU Data Boundary. For a UK contractor that answer needs checking, because the EU Data Boundary consists of the countries of the European Union and the European Free Trade Association, which Microsoft lists as the 27 EU member states plus Liechtenstein, Iceland, Norway and Switzerland 9.

The United Kingdom is not on that list. Microsoft's Copilot documentation says that EU traffic stays within the EU Data Boundary while worldwide traffic can be sent to the EU and other countries or regions for processing, and that customers outside the EU may have queries processed in the US, EU or other regions 7. A UK business assuring its client that its data stays inside the EU Data Boundary is, unless something specific has been arranged, very likely saying something untrue.

Two related details from the same documentation are worth knowing before you rely on a blanket assurance. The EU Data Boundary does not apply to web search queries: Microsoft states that Bing operates separately from Microsoft 365, is governed by the consumer Microsoft Services Agreement, and that Microsoft acts as an independent data controller for those queries 8. And the subprocessor picture is moving. As of 24 July 2026, OpenAI operated models are enabled for all users for eligible commercial customers unless an administrator turns them off in the admin centre, and for those models Microsoft records that there is no SOC 1 Type 2 report, no PCI DSS attestation, no HITRUST certification, and that they are currently excluded from in-country processing commitments where those apply 10. The core data protection terms still apply 10. The assurance stack around them is thinner, and the setting is recent enough to be worth checking rather than assuming.

Two different problems, and only one of them has a regulator

This is the distinction that makes a policy workable, and almost nothing published on this subject draws it.

Where the document contains personal data, UK data protection law applies and it is specific. Sending personal information to an organisation outside the UK is a restricted transfer, and the ICO states that the rules apply to all restricted transfers, even small, infrequent ones 14. A restricted transfer needs UK adequacy regulations, appropriate safeguards such as the International Data Transfer Agreement or the Addendum together with a transfer risk assessment, or an Article 49 exception, and where no appropriate exception can be identified the transfer must not be made 14. A site diary with operatives' names in it, a timesheet, a photograph with faces: these are the ordinary construction documents that engage all of that.

The commercial accounts are built for this and the consumer ones are not, which is the practical reason the tier distinction matters legally as well as commercially. Anthropic states that for its commercial products the customer is the controller of the data its users submit and Anthropic acts as a processor on the customer's behalf 6, and Microsoft states that organisational use of Copilot is covered by its Data Protection Addendum with Microsoft acting as a data processor 8. That allocation is what a written contract between controller and processor is supposed to record, and it is exactly what a personal account signed up with a private email address does not give you.

Where the document contains no personal data at all, and a great many construction documents do not, the exposure is contractual rather than regulatory. A drawing, a bill of quantities, a client's confidential tender information, a pricing schedule: the constraint on those is your non-disclosure agreements, the confidentiality provisions in the contract you signed, and any information security schedule your client imposed. No data protection assessment will tell you the answer, because no personal data is involved. One qualification matters: where the work is regulated surveying work, the RICS standard on responsible use of artificial intelligence applies to AI outputs with a material impact on that service regardless of whether personal data is present 17.

That second category is a genuine gap in the published guidance. No UK official guidance on putting commercially sensitive but non-personal information into AI tools was found while preparing this guide on 7 August 2026, which is a statement about what was searched for rather than a guarantee that none exists. The nearest official touchpoint is the National Cyber Security Centre's advice not to include sensitive information in queries to public large language models and not to submit anything that would cause a problem if made public 16, and that post dates from 2023 and describes public services rather than the commercial tiers a business would now buy. For most commercially sensitive material, your contract governs it, and the first place to look is the confidentiality clause and any client information security schedule rather than the ICO.

A policy that fits on one page

The useful version of this is short, because a long one does not get read and an unread policy protects nobody.

Name the approved tools, by account type. "Claude and ChatGPT" is not a policy. "The company Claude Team account and Microsoft 365 Copilot, and no personal accounts for company work" is one. Buy the commercial tier before you write the rule, because a rule that forbids what people are already doing, without giving them a permitted route, produces concealment rather than compliance.

Give two tests rather than a taxonomy. First: would you be comfortable if this document appeared in public, which is the NCSC's own framing 16 and needs no technical knowledge. Second: does this contain anyone's personal information, in which case it goes nowhere without checking the transfer position 14.

Then add the three things people do not think of. Turn off or explain the feedback buttons 4. Check who can see the Microsoft admin setting for subprocessor models 10. And write down what your client's contract actually says about confidentiality before you tell them anything reassuring about the EU.

Where a pilot is being run rather than a tool rolled out, the governance around it is covered separately in how to run a safe 30-day construction AI pilot, and the question of what Microsoft 365 Copilot can actually reach is in Copilot or a custom construction assistant.

What this guide cannot tell you, and what will change

Four limits, and they are unusually important on this subject.

It cannot tell you whether your specific use is lawful. That depends on your documents, your contracts, your transfer mechanism and your assessment, and this is an explanation of general principles rather than legal advice.

Several figures could not be verified and are therefore absent. Retention periods for ChatGPT Enterprise and Business were not established from primary documentation. Anthropic's commercial retention periods were not established. The UK to US data bridge position was not examined, so no conclusion is offered about the transfer mechanism for any named vendor.

The ICO's own AI guidance carries a notice that it is under review following the Data (Use and Access) Act, so the material a reader might expect to be settled is itself in motion 15. The international transfers guidance, updated in January 2026, is the more stable of the two 14.

And this page has a shorter shelf life than anything else in this library. Everything above was read on 7 August 2026. The Microsoft subprocessor position changed on 24 July 2026 and the documentation was edited on 5 August 2026 10, vendor product names are drifting, and the exclusion of one provider's models from the EU Data Boundary is described as current rather than permanent 7. Treat every specific claim here as needing a check, and treat the two anchor points, Anthropic's contractual clause 5 and Google's consumer warning 13, as the parts most likely to still be true when you read this.

The shorter version of this argument, framed around the three tiers rather than around the vendors, is on the blog as private AI versus public chatbots. Setting up a tenant so that the sanctioned route is also the easy one is the substance of our SharePoint and CDE work.

Sources

  1. 1.OpenAI, How your data is used to improve model performanceAuthorityAccessed
  2. 2.OpenAI, Data Controls FAQAuthorityAccessed
  3. 3.OpenAI, Data controls in the OpenAI platformAuthorityAccessed
  4. 4.Anthropic, Is my data used for model training? (commercial products)AuthorityAccessed
  5. 5.Anthropic, Commercial Terms of ServiceAuthorityAccessed
  6. 6.Anthropic, Does Anthropic act as a data processor or controller?AuthorityAccessed
  7. 7.Microsoft Learn, Data, Privacy, and Security for Microsoft 365 CopilotAuthorityAccessed
  8. 8.Microsoft Learn, Enterprise data protection in Microsoft 365 CopilotAuthorityAccessed
  9. 9.Microsoft Learn, What is the EU Data Boundary?AuthorityAccessed
  10. 10.Microsoft Learn, OpenAI as a subprocessor in Microsoft Online ServicesAuthorityAccessed
  11. 11.Microsoft Learn, Data, privacy, and security for Foundry Models sold by AzureAuthorityAccessed
  12. 12.Google, Generative AI in Google Workspace Privacy HubAuthorityAccessed
  13. 13.Google, Gemini Apps Privacy HubAuthorityAccessed
  14. 14.Information Commissioner's Office, A brief guide to international transfersPrimaryAccessed
  15. 15.Information Commissioner's Office, What are the accountability and governance implications of AI?PrimaryAccessed
  16. 16.National Cyber Security Centre, ChatGPT and large language models: what's the risk?AuthorityAccessed
  17. 17.Royal Institution of Chartered Surveyors, Responsible use of artificial intelligence in surveying practice, 1st editionAuthorityAccessed

Published , last reviewed . This guide explains general principles and is not legal, contractual or safety advice. The position on any project depends on the contract signed and the facts of that project.

If this is a problem you are carrying on a live package and you want to talk about what fixing it would take, get in touch or book a call.