AI governance
AI observability can't see a stale source
A retrieval trace records which document your AI reached for and how relevant the retriever thought it was. It carries no approver, no version and no effective date, so a faithfully grounded answer from a withdrawn policy still scores well.
Cognatum Team · Oct 1, 2026 · 5 min read
Your assistant answered a question about a retention policy. The trace is clean. You can see the query it ran. You can see the three documents it pulled back, and the score each one got. You can see the tokens it spent and the 1.4 seconds it took. None of that tells you the policy it quoted was replaced in March.
Cognatum governs the entry
source · version · approver · permissions
What the telemetry actually holds
AI observability is now a practice with its own vendors and its own words. Datadog describes it as measuring how models, data and responses behave. The point is to answer more than whether a system is running. It is to know whether the output is correct, grounded, safe and useful. Dynatrace maps the work across six layers. They run from the chat window down to GPU memory. The vector store and the retrieval step sit in the middle.
The shared standard for all of it is OpenTelemetry's GenAI semantic conventions. It is worth reading what a retrieval span is asked to carry.
Two fields per document
The attribute is gen_ai.retrieval.documents. The specification says each document object should contain at least two properties: an id, which is a unique identifier, and a score, which is the relevance the retriever assigned it. Everything else is optional or absent.
So the durable record of the material behind your answer is an identifier and a number. There is no field for the approver. None for the effective date. None for the version that was retrieved, or for whether a newer version exists, or for whether a second document in the same index says the opposite.
Grounded is not the same as right
Groundedness scoring is the part of observability that comes closest to the problem, and it is precise about its own scope. Datadog describes quality telemetry as evaluating a response against the provided context rather than against internal, possibly outdated or fabricated knowledge. That is the correct measure for the question it asks. It catches a model inventing a clause that is not there.
It cannot catch a model quoting a clause correctly from a document that was withdrawn last quarter. That answer is faithful to its context. The context was wrong. The score will be high.
The evidence on grounding alone
The first preregistered evaluation of commercial legal research tools measured this directly. The products tested were built on retrieval, and their vendors had described them as eliminating or avoiding hallucinations. The study found that the tools from LexisNexis and Thomson Reuters each hallucinated between 17 and 33 percent of the time. Retrieval lowered the rate against a general chatbot. It did not remove it.
When two sources disagree
Researchers surveying knowledge conflicts in large language models name three kinds: context-memory, inter-context and intra-memory. Inter-context conflict is the ordinary enterprise case. Two documents come back and they contradict each other.
A trace will record both. It will record both scores. A relevance score ranks how well a passage matches a query, which is a different question from which passage is true. Nothing in the telemetry marks that a contradiction was present, and nothing routes it to a person who could settle it.
Why this is not a missing feature
It would be easy to read this as a gap some vendor will close next year. It is more structural than that. Observability instruments the system. It records what the pipeline did, because that is what the pipeline knows.
Whether a document was approved, by whom, on what date, and whether it has since been superseded are not facts about the pipeline. They are facts about the document's standing in the organization. They live in whatever governs the document. If nothing governs it, there is nothing for the pipeline to emit.
You cannot instrument a field that does not exist
This is the uncomfortable part for a regulated team. The instrumentation is working exactly as designed. The record is thin because the knowledge underneath it was never given a version, an owner or an expiry date in the first place.
What the record is for
Regulators are asking for logs. Article 12 of the EU AI Act obliges high-risk systems to technically allow the automatic recording of events over the lifetime of the system. The official Service Desk page now carries a notice that the article has been amended and the displayed text is not yet updated, so check the current wording before relying on the detail.
The NIST AI Risk Management Framework treats documentation and traceability as part of what makes a system trustworthy, rather than as paperwork bolted on afterward. Neither instrument makes anyone compliant, and no software product confers compliance on its own. What software can do is supply the evidence a reviewer asks for.
Governing the knowledge, not just the call
This is the work the Cognatum Knowledge Loop is built around. Eight steps across four phases. AI runs seven of them: the finding, the gathering, the structuring, the surfacing. Approve is the one it does not run. An item is routed to a named person, who can rework it before signing off.
What that produces is a knowledge entry with an approver, a date and a source attached to it. The answer then inherits all three. Cognatum also keeps a record of where the knowledge stood at the moment the answer was given, which is point-in-time provenance rather than a current snapshot.
Conflicts and changes
Where two documents disagree, the system presents both and says where the disagreement is. It does not pick a winner, because that is a judgment call and judgment calls go to people. When a source changes, including a table, a query or a data mart rather than only a file, it notifies you that entries tied to it are now out of date. It does not quietly rewrite them, and it is not always current.
None of this requires moving anything. Knowledge stays in the departmental repositories where it already lives, and the flow runs in both directions.
The question telemetry cannot answer
Your traces can tell you, in detail, what your AI did. They can show which document it reached for and how confident the retriever was. They cannot tell you whether it should have reached for that one.
That is not a telemetry problem and no amount of instrumentation will fix it. Your company's knowledge isn't missing. It's unusable. Cognatum changes that.
See how the eight steps and the single human gate fit together in the Cognatum Knowledge Loop.
Sources
- What Is AI Observability? (datadoghq.com)
- What is AI observability? (dynatrace.com)
- GenAI semantic convention attributes (opentelemetry.io)
- LLM Observability: The 8 Best Tools for Production AI Systems (newrelic.com)
- Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools (arxiv.org)
- Knowledge Conflicts for LLMs: A Survey (arxiv.org)
- AI Act Article 12: Record-keeping (europa.eu)
- Artificial Intelligence Risk Management Framework, AI RMF 1.0 (nist.gov)
- The Cognatum Knowledge Loop (cognatum.ai)
Common questions
Questions this raises.
Does AI observability tell me whether an AI answer was correct?
Not on its own. Observability measures whether a response is faithful to the context it was given. If the retrieved document was out of date, or had never been approved, a faithful answer to it still scores well. Correctness depends on the standing of the source, and the telemetry does not carry that.
What does a retrieval trace actually record about a document?
Under OpenTelemetry's GenAI semantic conventions, each retrieved document object is expected to carry at least an id and a relevance score. There is no standard field for the version, the approver, the effective date or the retirement status of that document.
Is a high groundedness score enough for a regulated use case?
It is useful but not sufficient. The first preregistered study of retrieval-based legal research tools found hallucination rates between 17 and 33 percent despite vendor claims to the contrary. Grounding reduces invention. It does not establish that the grounding material was right.
What happens when two retrieved documents contradict each other?
Researchers call this inter-context conflict. A trace records both documents and both relevance scores, but a relevance score ranks match quality rather than truth. Cognatum surfaces the disagreement, says where it sits, and routes it to a person instead of picking a winner.
Does Cognatum keep itself up to date automatically?
No. It watches the sources it is pointed at, including structured sources such as tables, queries and data marts, and notifies you when entries tied to a changed source are now out of date. A person decides what the answer should be and approves it.