Cognatum

AI governance

Shadow AI: a knowledge problem before a security problem

Every definition of shadow AI describes the same behavior and prescribes the same fix: find the unapproved tools and block them. That treats the supply of AI tools as the problem. The demand is what actually needs answering.

Sep 9, 2026 · 6 min read

Search for shadow AI and the first page of results is almost entirely security vendors. Palo Alto Networks, IBM, Zscaler, Check Point, Wiz, Orca Security. Their definitions agree nearly word for word: the use of AI tools without the approval or oversight of IT or security. Their remedies agree too. Discover the tools, monitor the traffic, apply data loss prevention, block what is not on the approved list.

Served to one approved entry
AI assistants & agents
Proposal tools
Internal search & chat
Customer portals
Compliance & audit

Cognatum governs the entry

source · version · approver · permissions

That framing is accurate as far as it goes. It is also incomplete in a way that costs money. Detection tells you that someone pasted a client policy into a consumer chatbot on Tuesday. It does not tell you why they did it. In most cases the reason is mundane rather than reckless. Someone had a question, the approved systems could not answer it in a usable form, and a chatbot could.

The numbers describe demand, not deviance

CybSafe and the National Cybersecurity Alliance surveyed 7,000 workers in late 2024 and found that 38% share sensitive work information with AI tools without their employer's knowledge. IBM, drawing on the same period of research, notes that adoption of generative AI applications by enterprise employees rose from 74% to 96% between 2023 and 2024, and that one in five UK companies reported data leakage arising from employees using generative AI.

Two ways to read the same numbers

Read those figures as a security story and you get a workforce that ignores policy. Read them as a knowledge story and you get something more useful. A very large number of people, over a very short period, concluded that an unsanctioned tool was a better route to an answer than the sanctioned one. That is a judgment about the quality of the internal alternative.

What a breach costs

IBM's 2026 Cost of a Data Breach research puts a figure on what follows. The Cost of a Data Breach report found that more than 20% of organizations reported a breach targeting AI models or applications, with the most common causes being weaknesses in surrounding systems rather than the models themselves. Breaches at financial services organizations averaged $6.3 million in the same IBM research.

Blocking acts on supply, not on demand

The standard remedy list is sensible and worth doing. Acceptable use policies, an approved tool catalog, network monitoring, and training all reduce exposure. But each acts on the availability of outside tools. None acts on the reason people went looking.

The Cloud Security Alliance, which is not a lenient voice on governance, arrives at the same conclusion in its own guidance. Shadow AI, it argues, grows where employees see a quick path to solve pressing problems, and the answer is not elimination but governance and culture: somebody signs it off, and the record says who approved it. Its recommended path includes providing genuinely useful sanctioned alternatives, including retrieval over internal knowledge sources, rather than relying on prohibition alone.

The promise a ban makes

Every ban carries an implicit promise: use the approved route instead. When the approved route returns fourteen documents, three of which contradict each other, with no indication of which is current or who signed it off, the promise is not kept. The workaround returns, quieter than before.

What reviewers are actually asking about

FINRA's 2026 Annual Regulatory Oversight Report, published in December 2025, contains a standalone section on generative AI that substantially expands on the previous year. It reiterates that FINRA's framework is technology neutral and that firms remain responsible for compliance when using these tools, and it flags rules relating to supervision, communications, recordkeeping and fair dealing. It defines hallucination as instances where the model generates information that is inaccurate or misleading, yet is presented as factual information, and observes that a model misinterpreting a regulatory requirement can lead to flawed downstream decisions.

The monitoring expectations are the instructive part. FINRA points to four things:

  • reviewing prompts, responses and outputs over time
  • maintaining prompt and output logs for accountability
  • tracking which model version was used and when
  • conducting validation with human review

Notice what most of that list has in common. Very little of it concerns the model. Most of it concerns the content the system answered from, and whether there is a record of it. Elsewhere, the EU AI Act's Article 4 obligation requires providers and deployers to take measures supporting AI literacy among staff and others operating these systems on their behalf. The NIST AI Risk Management Framework, voluntary by design and currently being revised under the White House AI Action Plan, is organized around governing and documenting rather than around any particular model architecture. None of these confer compliance on a piece of software. What they do is define the evidence a reviewer will ask to see.

The supply side is a knowledge base

If shadow AI is partly a demand problem, the durable remedy is to make the approved answer easier to get than the unapproved one. That is a knowledge base question rather than a network question.

The distinction matters more than it sounds. Search can find a file. Search cannot tell you which answer inside that file was approved, by whom, or whether it still holds. An organization can have every relevant document indexed and still be unable to say which sentence is current. That is the condition most enterprises are actually in. Your company's knowledge isn't missing. It just isn't usable.

A record, not a second copy

Cognatum is a governed knowledge management system for regulated enterprises. It holds the approved answer itself rather than a second copy of your files, as a record with a type, an owner, a state and a version, and it serves that record to people, applications, workflows and AI systems with a named approver, a timestamp and a source on every entry. Information stays where it already lives, in the departmental repositories your teams already use, and Cognatum reads from those systems rather than asking anyone to migrate first.

The gap queue and the stale flag

Two behaviors matter for the shadow AI case specifically. Every question asked that has no approved record is logged as a gap, ranked by how often it was asked, by people and by assistants alike. That turns the demand signal into a queue owned by a named person instead of an invisible workaround. And when an underlying source changes, every answer built on it is flagged and routed for re-check. The system notifies you that something has gone stale. It does not quietly correct itself, and it is not always current by fiat, because a governed answer that changes without a person approving the change is not governed.

Where this leaves the security program

Nothing here argues against detection, policy or monitoring. It argues that those controls are the ceiling of what an organization can achieve while the internal alternative stays weak, and that the ceiling is lower than most programs assume.

The practical test is straightforward. Take the questions your people most often ask an outside chatbot, and ask whether your own systems could return an answer with a source, an approver and a date attached. Where they can, shadow AI is a policy matter. Where they cannot, it is a knowledge matter wearing a security costume, and no amount of blocking will settle it.

Knowledge governed. Intelligence everywhere.

Common questions

Questions this raises.

Is shadow AI just shadow IT with a new name?

They overlap but they are not the same risk. Shadow IT concerns unapproved software generally. Shadow AI concerns tools that ingest whatever a person types into them and produce confident output that then influences decisions. IBM draws the distinction around the nature of the tool: shadow AI introduces concerns specific to data handling, model outputs and decision making, which unapproved storage or project management tools do not.

Does blocking consumer AI tools solve the problem?

It reduces one avenue of exposure and is worth doing. It does not address why people sought an outside tool. The Cloud Security Alliance's guidance is explicit that elimination is not the solution, and recommends pairing controls with sanctioned alternatives that draw on internal knowledge sources.

What will a regulator actually ask about our AI assistant?

Based on FINRA's 2026 report, expect questions about supervision and recordkeeping rather than model internals: what the tool was used for, what governance and testing sat around it, whether prompts and outputs were logged, which model version was in use and when, and where human review applied. Requirements vary by regulator and jurisdiction, so treat this as the shape of the inquiry rather than a checklist.

Can software make us compliant with the EU AI Act or NIST AI RMF?

No. The NIST AI Risk Management Framework is voluntary and there is no certification attached to it, and obligations under the EU AI Act sit with providers and deployers rather than with a vendor. Software can align with these frameworks and supply the evidence they call for, such as approvers, dates, sources and versions. It cannot confer compliance, and any vendor claiming otherwise is worth a second look.

Where do we start if our knowledge is spread across a dozen systems?

Start with the questions rather than the systems. Identify what people actually ask, check whether an approved and current answer exists for each, and treat the gaps as the work. Migration is usually the wrong first move, because the departmental split of repositories is normal and generally fine. The problem is that nothing sits above those repositories to say which answer is approved.

Knowledge governed. Intelligence everywhere.

See it on your own content, in your own environment.