"Does it make things up" is actually two separate questions: does it get things wrong, and what happens to your data while it answers. Here is the honest version of both.
Why "does it make things up" is the wrong starting question
Generic AI chatbots answer from their training data alone, which is why they sometimes produce plausible-sounding nonsense. Notion AI works differently: it first searches your actual workspace, finds real candidate pages, and only generates its answer from that material.
That changes what "wrong answers" actually look like in practice. The most common failure mode is not the model fabricating a fact. It's one of these three:
- The right page doesn't exist yet. If nobody wrote the answer down anywhere, Notion AI has nothing to retrieve, and it will either say so or produce an incomplete answer assembled from whatever comes closest.
- The right page exists but doesn't surface. If it's poorly titled, buried, or contradicted by three stale versions of the same document, the ranking step may not bring it to the top.
- The page is found but misread. Even with the right source in hand, a model can still mis-summarize or miss the nuance buried in paragraph four.
The most important lever: your workspace quality
The first two cases are workspace problems, not AI problems. That's worth sitting with: the most important factor in Notion AI's accuracy for your team is not which model Notion uses. It's how well-organised your workspace is.
A team with one clean, up-to-date pricing page gets a reliable answer. A team with four stale copies of that same page scattered across old projects gets AI-generated guesses, delivered with confidence, because the search layer simply cannot know which version is current if your workspace doesn't tell it.
What this means day to day
Treat an AI answer the way you'd treat a summary from a capable junior colleague: probably right, worth a glance at the source before repeating it in a client meeting or a financial decision. Notion AI links to its sources when it can. Use that link.
The thumbs-up/thumbs-down buttons matter too, though it's worth knowing what they actually do: your feedback doesn't re-train the model on your specific workspace. It's sent to Notion's team to improve the product, not to make the AI progressively smarter about your content over time.
Data retention: how long your information stays with AI providers
By default, the LLM providers behind Notion AI retain your data for a maximum of 30 days before deletion, for workspaces not on Enterprise. On the Enterprise plan, the default is zero data retention: those providers do not store your data at all, not even temporarily, once they've responded to your request.
A small number of AI features may require data-retention models to function; when that's the case, admins must explicitly enable it in workspace settings, and it's off by default otherwise.
Embeddings follow their own timeline: they're deleted within 60 days of the page or workspace being deleted. If you delete a page or workspace, you have a 30-day window to restore it; after that it's gone, including any AI-generated content and the embeddings tied to it.
Where Notion AI stands on compliance
Notion AI is not treated as a less-audited product bolted on the side. It's explicitly included in the scope of Notion's SOC 2 Type 2 report and ISO 27001 certification, the same independent audits that cover the rest of the platform.
For regulated industries specifically, Notion AI can support HIPAA compliance by using the zero-retention APIs from LLM providers, which means protected health information doesn't linger anywhere it shouldn't.
The honest bottom line
No AI system from any vendor is 100% reliable, and Notion doesn't claim otherwise. What Notion AI has going for it is a genuine retrieval architecture that ties answers to verifiable content, permission boundaries it cannot cross, and a data retention policy that is genuinely stricter than most competitors by default, not just on the most expensive plan.
The realistic posture: use it to save time on first drafts and initial research, verify anything that actually matters before acting on it, and if the answers are bad, check your workspace's information architecture before blaming the model.