Why – and how – to ensure privacy when using GenAI (it’s in the anonymising!)

Generative AI tools are not bastions of privacy and security. Here’s what you need to know about anonymising documents to work carefully and securely.

Picture the scene: you are working on a document for a prospective new client, which they want kept confidential. You also want to use generative AI to, say, summarise the main points, or analyse the business proposal contained in it.

Is it safe to do the GenAI thing? Not at all.

Keeping the document itself secure on your local machine is sort-of simple. If the document is uploaded to Google Drive, you make sure that it sits in a folder only you can see. And if it is in the folders on your PC, you make sure that that PC is as secure as it can be. (My PC and laptop are double password protected, and I don’t let anyone else use them.)

When it comes to working with such texts in Microsoft Office, I have a paid account and take Microsoft at their word that the documents I work on are secure. (Manage document privacy | Microsoft Support)

But there are good reasons to pause and think before feeding the text into a GenAI tool like ChatGPT or Claude. In short, if you want to do that, you should take out all the identifying details first. 

Here’s a step-by-step guide to the whole process. It’s long, so here’s a list of the sections – feel free to scroll and cherry-pick.

Why should you be cautious about what you enter into a GenAI tool?

I have an extensive blog post on the subject, written in June 2025, and there’s no reason to think things have changed since then. 

READ: Privacy and Generative AI – what you need to know – Safe Hands

Don’t want to click? Here’s the TL;DR on that blog post:

We’ve been more-or-less-voluntarily giving personal information to Big Tech for years. With the advent of consumer-facing GenAI, we’re not just giving information to Big Tech – we’re giving information to tools which are essentially black boxes. There’s a lot we don’t know about how these models work, and a lot we can’t predict. Caution is the sensible thing here.

The risks we face now include:

  • Bigger = riskier: AI models are trained on massive datasets. This vast volume increases the likelihood of data exposure or misuse.
  • The things we can’t see: Gen AI systems often lack transparency regarding what data is collected, processed, stored, and shared. 
  • “No delete button”: Large Language Models (LLMs) don’t have a “delete” button or a straightforward mechanism to “unlearn” specific information. Once data is entered into most AI systems, it is challenging to remove it.
  • The legal stuff: Sharing private data with generative AI tools can violate data protection laws like POPIA. This is an important thing to think about if you are a freelancer, for example: your client data should not be AI fodder.

 Basic ways to work as securely as you can

There is one thing you can do to ring-fence the information you put into a GenAI tool: change the settings. On all the big platforms, there’s a way to turn off the tool’s ability to “train” on your data, which means it won’t use the text you input to help produce better outputs. In almost all cases, this setting is cunningly disguised – it will be phrased in a way that asks you to stop helping other people.

Here’s a list of where you can find it in the big tools, in the web browser versions on a PC:

ChatGPT: Go to your profile in the bottom left corner, then Settings, then Data Controls – turn off the option that says : “Improve the model for everyone.”

Claude: Go to your profile in the bottom left corner, then Settings, then Privacy – turn off the option that says: “Help improve our AI models.”

Gemini: Hit the settings (gear icon) in the bottom left corner, then Activity (which says: “Keeping your activity lets you pick up chats where you left off at any time and helps improve Google services, including AI models. When this setting is off, Google still saves chats for 72 hours to respond to you and help keep Gemini safe.”). Use the button to turn it off. This does mean you’ll lose previous chats, which means I have kept it turned on, and am careful about what I enter into Gemini.

Is a paid account more secure and private than a free one?

Yes and no.

If the generative AI is a paid, enterprise version (usually inside a company or organisation), the tool is as secure and private as the organisation’s IT can make it, based on a contract with the AI provider.

But paying for an individual account does not give you more privacy, according to several sources I consulted. The best of those was this: AI Data Privacy 2026: The AI Privacy Trap, which says: “If you are paying $20/month for ChatGPT Plus, Claude Pro, or Gemini Advanced, you are still likely training their models by default.” (So do the turning off recommended above!)

Ask the other party for permission

It’s best practice to ask your client (or family member or friend) for permission before you upload docs into GenAI tools. This might not seem important but think about it this way: if you knew a friend was telling her book club about something you told her in confidence, would you be happy? Same, same for Gen AI tools – if people share things with you, it is basic human etiquette to ask them for permission before spreading those things further.

See the Safe Hands AI policy

Assess the kind of document

Common sense is important here – there’s a lot of data you can share with an AI tool without thinking about it twice. If documents are already in the public domain, on the internet (a PDF on a company website, a government law, a report by a consultancy), have at it. GenAI tools are already able to access this information.

And if it’s yours to share, and you’re happy that anyone in the world could know about it, all good. Think recipes, advice for help with a sore knee, all the things you are talking to people about anyway.

From a business point of view, it also okay to share anything that wouldn’t be a surprise to a competitor. I happily ask AI tools for advice about marketing my business; it is surely no secret to anyone that I would do that, and the details of my business are on my website.

On the opposite end of the spectrum are the things you would never share – and a Gen AI tool would be no exception: 

  • passwords and security information
  • company information that is not in the public domain: product prototypes, confidential meeting notes, private research or executive travel plans
  • pictures of your children.

Proprietary details (whether they belong to you, or the company you work for) should never be entered into a chatbot conversation.

The middle ground: where the anonymising comes in

An example: you might want to ask an AI tool to assess your CV against a job advertisement. There’s no problem about the job ad – it is public information. But your CV might have your own full name, address, ID or social security number, passport, or driver’s licence details. And all of those could lay you open to identity theft.

So, the trick is to take out all the personal details before feeding the CV into the machine.

The same applies to bank statements: you want ChatGPT to help you do the household budget. Take out your name and bank account details first.

Same for medical information – a script from the pharmacy may have a wealth of information about you, or a family member. Remove!

Taking the details out – doc and PDF

So, what are the ground rules for anonymising?

This is completely stating the obvious but I’ll say it anyway: you shouldn’t feed the document into an AI tool and then ask it to remove the details. Doh!

The rule of thumb is to do the anonymising work away from AI tools that are cloud-based (meaning the data sits on servers that could be anywhere in the world). And that means doing some work. Here’s how to get the job done:

  1. Make a copy of the document and give it a non-identifiable filename: instead of Mike’s Dog Parlour Interim Financial Statements, call the file: Client A Interim Financial Statements. Then find all instances of the business name in the document and change them to something like “Client A”.
  2. If competitors or suppliers are mentioned, change those too.
  3. You might also have to change geographical details. If that dog parlour is the only such business in a small town, AI can pretty much figure out who it is. Change the area to a big city, or change the nature of the business.

To do all that, in Word and Drive, use the find/change functions. 

PDFs are more complicated; they aren’t easy to edit unless you are paying for a program that does that. A workaround: You can change a PDF to a Google Doc by uploading the PDF in Google Drive and opening it. At the top, to the right of the screen, you should see an option that says “Open With”. Click on that and open the PDF in Google Docs. (And see below in the section about local LLMs.)

Things to be aware of:

Using inbuilt tools like Gemini in Google Drive or Copilot in Microsoft to do the work could mean that the underlying Gen AI model is seeing the data – so don’t do that.

In Drive, if the document has tabs with details in, remember to remove those or change the information in all the tabs. AI will scan all tabs in a document.

Why you should try a local Large Language Model (LLM)

If that all seems like a lot of work, that’s because it is. This is a classic case of asking yourself whether using generative AI is worth all the trouble. Reading the documents and doing your own thinking and research might get the job done faster.

On the other hand, you can download a local LLM – a small Gen AI model which works only on your machine and is as safe as your machine is.

Background info: What is Local LLM?

I’ve written about this before (Gen AI that isn’t big tech – an exploration) and was, at that point, a bit underwhelmed by local LLMs. 

But it recently occurred to me that the killer use of a local LLM might be anonymising documents. I tested it, it works, and this is how I will be doing this task in future. 

How to get yourself a local LLM

I’ve done a short step-by-step below, but I’d recommend getting the specs of your computer and doing a bit of research about which might be your best bet.

1. My journey started when I saw this article: Gemma 4 sees, hears, and reads on my 16GB laptop, and it never phones home

2. I liked what I read and asked Gemini how to get Gemma 4, based on my system specs. It recommended a site called LM Studio (which is free for home and work use). LM Studio opened with a download for Windows, but its documentation says it works on newer Macs. (Read LM Studio’s system requirements here: System Requirements | LM Studio )

3. You download it, and do the usual opening of a .exe file.

4. Off the bat, LM Studio suggested I get something called “gemma-4-e4b” (which Gemini had described thusly: “Perfect / Blazing Fast. These ‘Effective’ edge-optimized models will run smoothly with incredibly fast response times (token-generation speeds) on your hardware.”

5. LM Studio launched, and after some back and forth about downloading the tool, I had a familiar-looking chat interface, with a field to enter my requests right at the bottom.

Screenshot of LM Studio interface

6. I tested it on a Safe Hands business proposal, in PDF format. All I asked it to do was to “take all identifying details out of this document”, thus testing it’s ability to deal with a vague request. Which it did, perfectly. It changed all instances of the client’s name to “Client Organization”, with other parties similarly disguised. The process was long (by AI standards); it took 12 minutes. Not the “blazing fast” promised by Gemini, of course. But on the other hand, it had worked seamlessly with a PDF. (And the speed, of course, relies on my underlying hardware; it might be that different models would work more speedily.)

Conclusion: I will be keeping Gemma 4, and I will be using it to anonymise documents in future.

Summary: the common sense rules

  • Be selective about what you share online (everywhere online!).
  • Read the AI tool’s privacy and security policies: Try to understand how they guard and protect your information.
  • Opt-out of the “let us train on your data” nonsense.
  • If working with client documentation is something you do, and you want to use generative AI in that process, investigate getting yourself a local LLM.

Final word: If you wouldn’t want it repeated, reviewed, or resurfaced later, it doesn’t belong in a cloud-based, individual (whether free or paid)chatbot inquiry.

Main picture:  Geralt, Pixabay

Previous Sensible Guide articles

Gen AI that isn’t big tech – an exploration – Large language model, small language model – what’s the difference? A step away from big tech, that’s what. 

Privacy and Generative AI – what you need to know – I’ve done some research on the question of privacy and Generative AI. My findings and thoughts…

Let’s go back to Gen AI basics – a list of basic GenAI concepts…

If any of this is relevant to your work…

You might find one of these useful:

Comments are closed.