Insytful AI Search security and data protection overview
Log in to add to favouritesPage last updated 07 October 2026
About this document
Insytful AI Search is a hosted service from Zengenti that answers website visitors' questions using the customer's own published web pages. This document is for IT security, information governance and data protection teams who are assessing the service.
It covers how the service works, what data it handles, where that data is stored and processed, and the controls around it. We publish it so you can get the answers you need without sending us a questionnaire first. If your process still needs a questionnaire or a DPIA, we're happy to complete it, and most of the answers will come from here.
Copies of our certificates and supporting evidence are available on request.
How AI Search works
Insytful AI Search uses retrieval-augmented generation (RAG). Zengenti doesn't build, train or fine-tune its own language model, so there's no training dataset and nothing is learned from your content or your visitors' questions.
- Indexing. Our crawler scans the customer's published website on a schedule. Pages are split into passages and converted into numerical representations (embeddings) using an open embedding model that runs on our own infrastructure. These are stored in a search index that belongs to that customer.
- Question. A visitor types a question. A safety check running on our infrastructure screens it for malicious or manipulative input. OpenAI's API is used to classify the question, check whether it may relate to a safeguarding concern, and rewrite it into a clearer search query. The rewritten query is turned into an embedding on our infrastructure and used to search that customer's index for the most relevant passages.
- Answer. The question and the handful of passages that best match it are sent to OpenAI's API, which writes a plain-language answer using only those passages. The answer is generated with a low randomness setting and shown with links to the source pages.
If the visitor asks a follow-up question, the earlier questions and answers in the same conversation are sent to OpenAI with it, so the follow-up can be understood in context.
The question goes to OpenAI in steps 2 and 3, along with any earlier turns in the same conversation and, in step 3, the matched passages. Everything else stays on our infrastructure, and OpenAI never receives the whole website or the search index.

What data is processed
The service draws only on content the customer has already published on its public website. It isn't designed to take in personal or special category data as a source of knowledge, and visitors don't need an account or have to give their name, email address or any other identifying details to use it.
The one place personal data can appear is in what visitors type. We don't ask for it, but a visitor could type their own name, a student or reference number, or details about their health into the search box or a feedback comment. We treat questions and comments as potentially containing personal data for that reason, which is why they're only kept for a limited period.
| Data | Where it comes from | Could it contain personal data? | Kept for |
|---|---|---|---|
| Website content | The customer's published pages | Only what the customer has already published | Until the page changes or is removed and the site is re-indexed |
| Visitor questions and answers | Typed by visitors, generated by the service | Possibly, if a visitor types it in | 2 weeks by default, then deleted. Each customer can set a different period, up to a maximum of 90 days |
| Conversation history | Earlier questions and answers in the same conversation, so follow-up questions work | Possibly, if a visitor types it in | 7 days, then deleted |
| Answer feedback ratings | Thumbs up or down given by visitors on an answer | No | Retained for reporting |
| Answer feedback comments | Optional free text typed by visitors with a rating | Possibly, if a visitor types it in | Deleted at the same time as the answer they refer to |
| Aggregate usage metrics | Derived from search activity | No | Retained for reporting |
| Data sent to OpenAI | The question, earlier turns in the same conversation, and the matched passages | Possibly, if a visitor types it in | Up to 30 days at OpenAI, see sub-processors |
Where data is stored and processed
Everything except question classification, query rewriting and answer composition runs on Zengenti-operated infrastructure in UK Tier 3, ISO 27001-certified data centres, with colocation sites in London and Manchester and a further site in Ludlow. Stored data is held in Contensis and in databases we run on that same infrastructure.
| Part of the service | Location |
|---|---|
| Crawler, search index and embedding model | UK, Zengenti infrastructure |
| Safety check on incoming questions | UK, Zengenti infrastructure |
| Question and answer logs, conversation history and feedback | UK, Zengenti infrastructure |
| Analytics and reporting | UK, Zengenti infrastructure |
| Question classification, query rewriting and answer composition | OpenAI API, which may process data outside the UK, including in the US |
We use OpenAI's standard API without a regional processing setting, so we can't guarantee where OpenAI processes each request. What OpenAI receives is limited to the visitor's question, any earlier turns in the same conversation, and the passages of public website content needed to answer it.
Sub-processors
OpenAI is the only sub-processor that receives data from Insytful AI Search.
| OpenAI | |
|---|---|
| Purpose | Classifying each question, including whether it may relate to a safeguarding concern, rewriting it as a search query, and composing the wording of each answer |
| Data received | The visitor's question, earlier turns in the same conversation, and the matched passages of public website content |
| Location | May be outside the UK, including the US |
| Retention | Up to 30 days in OpenAI's abuse monitoring logs, then deleted, unless longer retention is required by law |
| Used for training | No. OpenAI doesn't use API data to train its models unless the customer opts in, and we haven't |
| Model | OpenAI's language models. We review which model we use and may change it over time |
OpenAI's own terms are set out in its API data controls documentation. We'll tell customers in advance if we add or change a sub-processor.
Retention, deletion and use of data
Visitor questions and the answers returned are deleted after two weeks by default. Each customer can choose a different retention period, up to a maximum of 90 days. Conversation history used for follow-up questions is deleted after 7 days. Feedback comments are deleted at the same time as the answer they refer to, while the thumbs up or down rating itself is kept for reporting because it contains no personal data. Aggregate metrics derived from search activity, such as question volumes and how often a page is cited, are kept for reporting and don't include personal data.
No data from the service is used to train an AI model, either by Zengenti or by OpenAI. To check the service is working well, our team reviews individual questions and answers within the retention period to test answer quality and spot anomalies. That's done by people, not by feeding the questions into a model, and the questions are deleted on the same schedule as everything else.
Website content stays in the index while it's live on the customer's site. When a page is changed or removed, the index reflects that at the next re-index, and customers can exclude individual pages or whole sections from AI Search at any time.
Deleted data can remain in backups until the customer's backup retention period runs out, after which it's overwritten.
Security controls and certifications
Zengenti holds ISO 27001:2022, ISO 9001 and Cyber Essentials. The ISO 27001 scope covers all company operations, including software development, hosting, support and maintenance, and it's externally audited every year. Our security policies are reviewed under that system and are available on request.
- Hosting and physical security. Our production infrastructure runs in two UK colocation data centres, Interxion in London and Equinix in Manchester, which are run by separate companies. Both are Uptime Institute Tier 3 and hold ISO 27001 in their own right. They have CCTV, 24/7 security guards and patrols, and access requires government-issued ID plus pre-approval from Zengenti.
- Penetration testing. The Insytful platform has an annual penetration test carried out by a certified third-party testing company. Customers are also welcome to commission their own tests.
- Encryption in transit. All traffic between the visitor's browser and our infrastructure, between components, and between our infrastructure and OpenAI uses TLS (HTTPS).
- Data at rest. Stored data, including the search index, question and answer logs, conversation history and feedback, is held in Contensis and in databases on our own UK infrastructure. It isn't encrypted at rest. It's protected instead by the physical, network and access controls described here, and visitor questions are deleted at the end of the retention period.
- Access control. Access to stored data and to the hosting platform is limited to authorised Zengenti staff who need it for their role. All staff complete data protection training when they join and a refresher every year.
- Tenant separation. Each customer has its own search index within a shared cluster. Access is scoped at the application layer, so a question can only ever be searched against the index for the site it was asked on.
- Network protection and patching. The platform sits behind firewalls with baseline DDoS protection. Every component has its own patching plan, and critical security patches are applied straight away through our incident management process. Servers run in resilient pairs, so patching doesn't take the service offline.
- Backups. Data is backed up automatically every night and replicated to a second UK data centre and to air-gapped offline storage. Each customer can set how long its backups are kept, and we test full restores every quarter.
- Monitoring. The platform is monitored 24/7/365 for availability, response times, capacity and errors, with automated alerts to on-call engineers and escalation if an alert isn't picked up.
- Secure development. Our software development follows OWASP principles, and every release is tested before it's deployed.
- Support. Support runs through an ISO 27001-certified service desk with defined response and resolution targets.
AI safeguards and answer quality
The main control against made-up or unsupported answers is the RAG design itself. The model isn't asked what it knows about a subject. It's asked to answer using passages retrieved from the customer's own site, and every answer links to the pages it came from so visitors can check the source.
- Incoming questions are screened by a safety check before processing, and prompt injection is one of the categories we actively test.
- Questions that may relate to a safeguarding concern are identified so the answer can point to the right help.
- Answers are generated with a low randomness setting.
- We run structured quality testing across realistic scenarios, including ambiguous and vague questions, multi-turn conversations, typing errors and prompt injection attempts.
- Because the service only draws on the customer's own published content, there's no outside data source for anyone to tamper with.
These measures reduce the risk of a wrong answer but don't remove it. If a page is out of date or contradicts another page, that can show up in answers. Customers stay in control of this, because they can correct or remove the source page, exclude pages or sections from AI Search, or have the service or individual features switched off.
Roles, responsibilities and DPIAs
The customer is the data controller for questions asked on its website. Zengenti is the data processor, and OpenAI is our sub-processor.
- Data processing agreement. Data processing agreement. We're happy to sign a data processing agreement with each customer, and we usually work from the customer's own template. It can cover AI Search alongside any other Zengenti services you use, and it will list OpenAI as our sub-processor.
- DPIAs. We've supported customers through their own DPIAs for AI Search and can provide answers, diagrams and supporting detail for yours.
- Privacy notices. We'd recommend mentioning AI Search in your website privacy notice, covering what visitors type in, how long it's kept and that answers are composed by a third-party AI model. We can suggest wording.
- Visitor guidance. You can add a short line near the search box asking people not to enter personal details, which reduces how often personal data reaches the service at all.
Incidents and contacts
Security incidents are handled under our ISO 27001 incident management process, with on-call incident managers who keep affected customers updated while an incident is being resolved. If an incident affects customer data, we'll notify the affected customers [timescale to confirm] and work with them on any reporting they need to do.
| Area | Contact |
|---|---|
| Data protection | Carl Gottlieb, Data Protection Officer |
| Information security | Rich Chivers, CEO, who holds CISO responsibilities |
| AI Search architecture | Derry Coffey, Lead AI Engineer |
| Everything else | Joe Miller, Product Manager |
Common questions
| Question | Short answer |
|---|---|
| Is our data used to train AI models? | No, not by Zengenti or by OpenAI. |
| Does OpenAI see our whole website? | No. It only receives each question, any earlier turns in the same conversation, and the few passages needed to answer it. |
| Where is our data stored? | In the UK, on Zengenti infrastructure. Question classification, query rewriting and answer composition by OpenAI may happen outside the UK. |
| Is data encrypted? | It's encrypted in transit using TLS. Stored data isn't encrypted at rest, but it's held in access-controlled UK data centres and visitor questions are deleted at the end of the retention period. |
| How long are visitors' questions kept? | Two weeks by default, then they're deleted. You can choose a different period, up to 90 days. |
| Do visitors need to log in or identify themselves? | No. The only thing a visitor provides is their question, and optionally feedback on an answer. |
| Can other customers' searches reach our content? | No. Each customer has its own index and questions are scoped to it. |
| Can we stop certain pages being used? | Yes. Pages and whole sections can be excluded. |
| Can AI Search be switched off? | Yes, for the whole service or for individual features. |
| Do you hold Cyber Essentials Plus? | We hold Cyber Essentials, ISO 27001:2022 and ISO 9001. |
Sources
- OpenAI API data controls, for OpenAI's training, retention and data residency terms