Your AI knowledge base can find the fact and still miss the decision
Retrieval tells you whether the system can access the information. Decision sensitivity tells you whether that information changes the judgement.
An enterprise AI knowledge base can retrieve the correct fact and still fail to use it in a decision. In a 2026 preprint, Miao Liu and Zhizhe Liu found that an AI analyst retrieved the relevant risk disclosure for all 12 firms at 128,000 tokens, while the disclosure's influence on its investment judgement fell to the experimental noise floor. A targeted, structured extraction beside the judgement restored that influence. The result does not prove that every RAG or long-context system fails in the same way. It shows that retrieval accuracy is an incomplete acceptance test. In the Brain Pillar of the Havruta Methodology™, Ground Truth is connected to decision logic through an Institutional Data Layer: facts, definitions, rules, exceptions, rationale, sources, owners, versions, confidence and gaps. Leaders should therefore test whether changing one material fact changes the recommendation, reasoning or confidence. If the judgement does not move, the system may have built search without building a business brain.
On this page
- The fact was found. The judgement did not move.
- The test most enterprise AI projects stop at
- What the 2026 study actually tested
- Retrieval succeeded. Decision use failed.
- What this means for enterprise AI
- Searchable library versus business brain
- The Brain Pillar and Institutional Data Layer
- The material-fact test
- Four ways to run the test
- Who owns this inside the organisation?
- What leaders should ask before scaling
- Frequently asked questions
- References
The fact was found. The judgement did not move.
Read the video transcript
Your AI can find the right fact and still ignore it. A recent study tested AI analysts on long financial filings. At 128,000 tokens, the primary model retrieved the risk disclosure for all 12 firms. But the disclosure no longer changed its judgement beyond experimental noise. The fact was available. It was not shaping the decision. A targeted, structured restatement beside the decision restored its influence. So test your system. Change one material fact. Ask again. If the judgement does not move, you built search, not a business brain.
The test most enterprise AI projects stop at
A contract system finds the clause. A finance assistant cites the forecast assumption. An HR copilot quotes the policy exception. A sales agent retrieves the customer concentration figure.
Each result looks like evidence that the knowledge system works.
It is evidence of something important. The system can access the source, locate the relevant passage, and return it in a useful form. Those are necessary tests. A knowledge system that cannot find the information has failed before any judgement begins.
The trouble starts when retrieval becomes the finish line.
Most acceptance tests ask whether the system can find the relevant document, cite it correctly, summarise it accurately, and respond consistently. None of those tests shows whether the fact changed the recommendation it was meant to inform.
Imagine that a contract contains a change-of-control clause that should alter the deal structure. The AI retrieves it and quotes it. The recommended structure stays the same.
Or a forecast assumes stable demand. The assumption changes materially, yet the plan, confidence, and capacity recommendation remain untouched.
The system did not miss the file. It missed the consequence.
That is a harder failure to see because the answer still looks grounded. It has citations. It uses the organisation's language. It may even explain the relevant passage when asked. Fluency and source access create confidence long before anyone has tested whether the information became active in the judgement.
For a searchable library, retrieval may be the outcome. For a system sold or governed as decision support, it is the entry condition.
If you need the foundations first, start with what an AI knowledge base is and how it works.
What the 2026 study actually tested
Miao Liu and Zhizhe Liu's preliminary paper, Reading Is Not Using, separates questions that enterprise AI evaluations often blend together.
- Was the fact present?
- Could the model retrieve it?
- Did it change the judgement?
The researchers created matched financial filings for 12 US registrants. In one version, they inserted a firm-specific risk disclosure, such as a covenant threshold, settlement payment, or indemnification cap. Control versions replaced that paragraph with equal-length neutral material.
The material disclosure stayed fixed. Only the unrelated surrounding content grew, from 2,000 to 128,000 tokens. Retrieval and investment judgement were then tested separately. That separation let the researchers observe whether the model could state the fact and whether the fact changed its probability of recommending a sale.
At 2,000 tokens, the inserted disclosure increased the primary model's sell probability by 3.2 percentage points. Between 8,000 and 32,000 tokens, its influence became statistically indistinguishable from the effect of inserting neutral text. It stayed at that experimental noise floor through 128,000 tokens.
Retrieval followed a different path. At 128,000 tokens, the same model retrieved the risk disclosure for all 12 firms and produced no false retrievals on the neutral filings.
The fact was present. The model could find it. The judgement did not move in a measurable way.
Text equivalent: stage one, the fact is present in the filing. Stage two, the fact is retrieved, for all 12 firms. Stage three, the judgement changes: at 128,000 tokens this stage failed, with influence at the experimental noise floor.
The study then tested several workflows. More reasoning did not restore the lost influence. A generic chunk-and-summarise approach performed worse because its notes dropped the target disclosure before the decision stage. Repeating the raw paragraph beside the decision did little.
A targeted, structured restatement placed immediately before the judgement changed the result. At 128,000 tokens, the disclosure shifted judgement by 8.5 percentage points, and all 12 firms moved in the predicted direction.
That result needs its boundaries beside it.
What the study does not establish
The paper is a preliminary preprint and has not completed peer review. Its main disclosures were written by the researchers. The primary task was financial analysis, not a deployed corporate workflow. A separate experiment on real 10-K disclosures pointed in the same direction, but the authors describe that evidence as exploratory and capability-dependent.
The study does not estimate how often this failure appears in enterprise systems. It does not establish one context-length threshold for every model. Stronger models moved the failure boundary in the tests, and the largest open-weight model retained measurable influence at 128,000 tokens.
Nor does the paper show that one simple formatting trick solves the problem. The successful intervention combined targeting, structure, compactness, fidelity, wording, and proximity to the decision. The study does not isolate one of those ingredients as the whole cause.
What it demonstrates is narrower and useful: retrieval accuracy and decision influence can separate.
Retrieval succeeded. Decision use failed.
Information can be available to a model without being active in its judgement.
Liu and Liu call this the retrieval-integration gap. The model forms a usable representation of the disclosure where it reads it. As competing context grows, the channels carrying that information into the final judgement weaken.
For a leader reviewing an enterprise AI system, the internal mechanism matters less than the management consequence. A quotation is not proof of weighting. A citation is not proof of influence.
The system may retrieve a clause after the recommendation has been formed. It may explain a policy exception when directly asked about the exception, yet fail to apply it when choosing between two organisational options. It may cite a forecast assumption in the evidence list while producing almost the same plan after that assumption changes.
That is why a good-looking answer can be deceptive. The source appears. The language is accurate. The reasoning reads smoothly. But the decision remains insensitive to a fact the accountable owner considers material.
There is a second trap. Teams often test the answer after drawing the model's attention to the exact fact they want checked. That proves the model can respond when the evaluator has already done the work of deciding what matters. It does not show that the system will carry the fact into a wider decision when it competes with dozens of plausible considerations.
The better test preserves that competition. Give the system the full evidence it is expected to use, ask for the real decision, and observe whether the material fact changes the result. Then run the counterfactual. The difference between the two judgements is the signal that retrieval tests leave unseen.
Retrieval accuracy is an incomplete acceptance test.
What this means for enterprise AI
The paper does not prove that every enterprise knowledge system fails this way. It gives leaders a reason to test for the failure rather than assume retrieval has solved it.
A repository holds records. Search makes those records accessible. Retrieval-augmented generation can place relevant passages into a model's working context. These are real advances. They reduce the time spent finding and retyping information.
They do not, by themselves, explain why the information matters.
A policy may contain an exception without stating which organisational goal the exception protects. A forecast may record an assumption without naming which downstream commitments should move when it changes. A contract may set a threshold without carrying the commercial priority that determines whether the organisation should accept, renegotiate, or walk away.
Documents preserve pieces of truth. Decisions require relationships between those pieces.
This is where many enterprise AI projects become technically impressive and managerially vague. The system is connected to more content, but the organisation has not defined which decisions should change, which facts are material to them, or what movement would count as correct.
That is the same value problem behind enterprise AI investment without a changed business outcome. Activity can rise while the decision remains untouched.
NIST's August 2026 draft TEVV-Athlon framework makes the wider evaluation point: tests must be tailored to the application and the real-world outcome being assessed. If the application's job is to support a decision, then the evaluation must reach the decision.
That changes the acceptance question.
Do not ask only, “Can the AI find the policy?” Ask, “When the policy changes, what should move?”
Searchable library versus business brain
A searchable AI knowledge base and an enterprise business brain are not opposing systems. The second depends on the first. The difference is the layer added after access.
| Dimension | Searchable AI knowledge base | Enterprise business brain |
|---|---|---|
| Primary question | What information can the system find? | What should change because of the information? |
| Core unit | Document or passage | Maintained brain containing facts and operating logic |
| Context | Source text | Goal, rationale, relationships, and consequences |
| Rules | Often implicit in documents | Explicit rules, thresholds, and exceptions |
| Ownership | Repository or platform owner | Named domain owner owns truth |
| Maintenance | File added or replaced | Version, confidence, gaps, and rationale maintained |
| Evaluation | Retrieval and answer accuracy | Decision sensitivity, reasoning quality, and value |
| Scale condition | More content connected | Proved value plus leadership pull |
The searchable library answers a location question. Where is the clause? Which policy applies? What did last quarter's plan say?
The business brain answers a consequence question. Given this clause, policy, or assumption, what changes in the decision, and why?
That requires goals, decision rules, thresholds, exceptions, owners, and escalation paths to become explicit. Much of that logic is not sitting neatly inside one document. It lives across policy, practice, history, and the judgement of people who know why the organisation works the way it does.
Connecting another folder will not create that logic. Someone has to surface it, test it, assign ownership, and keep it current.
The word “brain” matters because the object is maintained, not merely stored. When a rule changes, the reason for the change travels with it. When confidence is low, the gap remains visible instead of hardening into an answer. When two sources conflict, the conflict has an owner. A pile of current documents can still leave each of those questions unresolved.
This distinction also changes the economics of scale. Adding content to a repository is cheap. Maintaining decision logic is selective work. The organisation should begin with the decisions where a better judgement has a clear consequence, prove that the missing logic can be built and used, and scale only after the value is visible.
The Brain Pillar and Institutional Data Layer
The Havruta Methodology™ addresses this through its Brain Pillar: the persistent substrate that holds across dialogues.
The Brain Pillar starts from a simple constraint. If the machine has to guess the organisation's goals, definitions, relationships, and decision rules every time, the output may be fluent but it is not grounded in the way the business actually works.
Ground Truth is the leader's and organisation's domain knowledge, data, judgement, audience understanding, and contextual sensitivity. But truth alone is not enough if it remains scattered across files or trapped in someone's head. Ground Truth must be connected to the logic that tells the system when a decision should change.
That connected layer is the Institutional Data Layer. It contains facts, definitions, rules, exceptions, rationale, sources, owners, versions, confidence, and gaps. Its unit is the brain, not the uploaded file.
A brain is a maintained description of how a domain works. It can state that a threshold exists, why it exists, which goal it protects, when an exception applies, who may approve that exception, and what downstream plans should change if the threshold is crossed.
This also explains the distinction between an agent and a brain. An agent is the wrapper that performs a task. The brain is the substance it reasons from. A capable wrapper connected to a large document store may still lack the operating logic needed to handle a changed fact.
At organisational scale, the work needs clear roles.
AI Contributors are business-deep practitioners trained to lead the creation of institutional data and decision logic. They find a real bottleneck, reach the people who hold the missing knowledge, and turn that knowledge into a maintained brain. They do not become the owners of every fact they collect.
Domain owners retain ownership of truth. Leaders choose which decisions matter, protect the time needed to build, and create Leadership Pull by using the capability on real work. Engineering, data, security, legal, and governance teams productionise what has proved value.
This is the organisation-scale operating model of the Brain Pillar. It is not a third pillar and it is not a substitute for technical architecture. It connects business truth and decision logic to the systems that deliver them.
The material-fact test
The fastest way to expose the gap is to change one fact that should matter.
This is a practical management diagnostic derived from Liu and Liu's finding and the Havruta Methodology™. It is not a benchmark proposed or validated by the paper.
1. Choose a real decision
Use a decision with an accountable owner and a consequence. “What does the policy say?” is a retrieval question. “Should we approve this exception?” is a decision.
2. Name one material fact
Ask the domain owner to identify a fact that should change the decision. Agree on the expected direction before running the test. If nobody can say what should move, the decision logic is not yet clear enough to evaluate.
3. Record the baseline
Capture the recommendation, reasoning, confidence, evidence cited, material assumptions, and any escalation requested. A final answer alone is too thin. You need a record of how the system reached it.
4. Change only the material fact
Keep the goal, question, sources, and other conditions stable. Change the threshold, assumption, priority, or exception being tested. If several variables move at once, you will not know which one influenced the result.
5. Repeat the decision
Ask the same question under the changed fact. Keep the system, instruction, and evidence set as stable as the test allows.
6. Inspect the movement
Compare more than the final recommendation. Did the reasoning change? Did confidence rise or fall? Was a new risk surfaced? Did the escalation path change? Did the system cite the evidence that caused the movement?
Three verdicts
Moves as expected. The fact appears active in the decision. That is evidence, not final proof. Repeat the test across more cases, edge conditions, and owners before scaling.
Moves for the wrong reason. The system is sensitive, but the evidence chain is weak or the reasoning conflicts with the domain owner's logic. Repair the missing rule, source, or relationship, then test again.
Does not move. Retrieval may be working while decision integration is not. Stop the scale-up. Diagnose whether the system lacks the decision rule, the fact is poorly represented, the workflow separates evidence from judgement, or the supposed materiality was never agreed.
Copy this into your AI system
Act as the accountable decision owner for the decision described below. Use only the supplied organisational evidence. First state the recommendation, reasoning, confidence, material assumptions and evidence used. Then ask me which single fact I want to change. After I change it, repeat the analysis and explain exactly what moved, what did not and why. If the information is insufficient, stop and ask me one question at a time.
Which decision should your AI change?
Start with one real decision, one material fact, and one accountable owner. If the system cannot explain what changed and why, it is not ready to scale.
Four ways to run the test
These examples are synthetic. Each follows the same pattern: fact, decision logic, expected movement.
Legal: the priority behind the agreement changes
Decision: Prepare the negotiating position for a property sale.
Material fact: The organisation's priority changes from maximising price to completing before a financing deadline.
Expected movement: The AI should change the recommended negotiating position, the trade-offs it accepts, the sequence of work, and the clauses that need early resolution. A system that merely rewrites the same contract faster has retrieved the legal material without applying the changed commercial priority.
Finance: a core assumption falls
Decision: Approve next year's operating plan.
Material fact: The core volume assumption falls, while a dependent capacity investment remains fixed.
Expected movement: The AI should revise the forecast, identify the plans tied to the old volume assumption, and change the recommendation or escalation path. If the plan remains stable, the assumption may be visible in the model without being connected to its consequences.
HR: the assumed capability is absent
Decision: Recommend an organisational-design option.
Material fact: A capability assumed to exist internally is shown to be absent.
Expected movement: The AI should change the build, buy, partner, or sequencing recommendation. Repeating the preferred organisation chart with a paragraph about the capability gap is not enough. The missing capability must alter the route.
Sales: the account falls below the threshold
Decision: Decide whether to pursue a major account.
Material fact: Buying authority is weaker than expected and the required margin falls below the approved threshold.
Expected movement: The AI should change qualification, investment, or escalation. A better account plan for an account that no longer meets the organisation's rules is a polished failure.
Who owns this inside the organisation?
No single function holds the whole system.
Leaders choose the decisions worth testing. They make the priority visible, remove barriers, and decide whether the evidence justifies further investment.
Domain owners define what is true. They decide which facts are material, what movement is expected, which exceptions are legitimate, and where the system must stop and ask for judgement.
AI Contributors lead the work of eliciting and structuring the missing logic. They connect sources, rules, rationale, ownership, confidence, and gaps into maintained brains that can be tested.
Engineering and data teams connect those brains to the right systems, control access, and make the capability reliable. Governance defines permissions, traceability, audit requirements, escalation, and stop conditions.
The operating problem is coordination, not competence. Business truth without production engineering remains a useful file. Engineering without explicit decision logic produces a capable system that may still guess what matters. Governance without a tested decision can control activity without knowing whether the capability creates value.
Maintenance is part of ownership. A decision test that passes today can fail six months later because the policy changed, the source was replaced, the threshold moved, or the person who understood the exception left. The brain needs a named review rhythm and a clear record of what changed. Otherwise yesterday's logic becomes today's confident error.
Leaders should resist the urge to solve this by forming a large committee. Start narrow. One decision exposes which people, sources, rules, and controls are actually needed. That small build gives every function something concrete to inspect before a programme is expanded.
The work meets at one scale gate: has the capability proved value on a real decision, and do leaders understand it well enough to create pull for responsible use?
What leaders should ask before scaling
Before connecting more documents, adding more users, or moving the system into a higher-consequence workflow, ask five questions.
- Which decisions should this system change? Name the decision, the owner, and the consequence.
- Which facts should materially alter those decisions? Agree on expected movement before testing the model.
- Is the decision logic explicit, current, and owned? A rule without an owner will drift.
- Can we trace why the judgement moved or did not move? Citation alone does not answer that question.
- What evidence would justify scaling this capability? Define the bar before enthusiasm lowers it.
A searchable library answers what the organisation knows. A business brain must also show what should change because of it.
Frequently asked questions
Can an AI retrieve the correct information and still fail to use it?
Yes. Liu and Liu's 2026 preprint found that the primary AI model retrieved the relevant risk disclosure for all 12 firms at 128,000 tokens, while the disclosure's influence on its investment judgement had fallen to the experimental noise floor. The study calls this a retrieval-integration gap. It does not show that every AI system fails in the same way. It shows why retrieval and decision influence need separate tests.
Is retrieval accuracy enough to evaluate a RAG system?
No, if the RAG system is meant to support a decision. Retrieval accuracy tells you whether the system found the relevant evidence. It does not tell you whether that evidence changed the recommendation, reasoning, confidence, risk treatment, or escalation path. Retrieval remains necessary. The next test is decision sensitivity: change one material fact, repeat the decision, and inspect whether the judgement moves in the expected direction.
What is the difference between an AI knowledge base and a business brain?
An AI knowledge base makes documents and passages accessible. A business brain adds the operating logic that explains what should change because of the information. That includes goals, rules, thresholds, exceptions, rationale, ownership, and consequences. The business brain uses retrieval rather than replacing it. Its higher standard is whether current Ground Truth changes the decision in a way the accountable domain owner can inspect and defend. The Agent vs Brain distinction explains why the wrapper alone cannot supply that substance.
What is decision logic in enterprise AI?
Decision logic is the maintained set of rules and relationships connecting a fact to a consequence. It explains why a threshold matters, which goal it affects, what exception applies, who owns the truth, and when escalation is required. In the Brain Pillar, this logic sits beside the source material so the AI system does not have to infer the organisation's priorities from documents written for another purpose.
How do you test whether AI knowledge changes a decision?
Use the material-fact test. Choose a real decision, agree with the domain owner on one fact that should change it, and record the baseline recommendation and reasoning. Change only that fact, repeat the same decision, then compare the recommendation, confidence, evidence, risk treatment, and escalation path. If nothing meaningful moves, stop the scale-up and diagnose the missing decision rule or information flow.
What belongs in an Institutional Data Layer?
An Institutional Data Layer contains facts, definitions, rules, exceptions, rationale, sources, owners, versions, confidence, and gaps. It is broader than a technical data platform. A warehouse can store records. The Institutional Data Layer also carries the operating logic needed to reason about a change. Its unit is a maintained brain, not an uploaded file, and each domain owner remains accountable for what is true.
What is the Brain Pillar of the Havruta Methodology™?
The Brain Pillar is the persistent substrate that holds across AI dialogues. It carries the organisation's established facts, definitions, relationships, and decision rules so each new interaction does not begin from zero. It is one of two pillars in the Havruta Methodology™. The Cognitive Pillar shapes the live dialogue. The Brain Pillar preserves what the AI partner must know accurately between dialogues.
Who owns the truth in an enterprise AI system?
The domain owner owns the truth. An AI Contributor may lead the elicitation, structure, source discipline, gap management, and maintenance needed to build the Institutional Data Layer, but that does not transfer accountability for the content. Engineering owns production architecture and integration. Governance owns permissions, auditability, and stop conditions. Clear ownership matters because a current source without an accountable validator can become stale authority.
What should happen if the AI's judgement does not change?
Stop scaling that decision workflow and diagnose the cause. Check whether the changed fact was truly material, whether the decision rule was explicit, whether the evidence reached the point of judgement, and whether the system can explain its weighting. Do not treat a correct citation as a pass. A judgement that does not move may show that retrieval works while integration does not, which is exactly the gap the test is designed to expose.
Does this research prove that long-context AI is unreliable?
No. Reading Is Not Using is a preliminary financial-analysis preprint, not a universal reliability study. The tested models differed, and stronger capability moved the point at which decision influence weakened. The main disclosures were researcher-authored, while the real-filing experiment was exploratory. The safe conclusion is narrower: advertised context length and accurate retrieval do not, on their own, prove that material information shaped the judgement.
References
- Miao Liu and Zhizhe Liu, Reading Is Not Using: Retrieval, Judgment, and the Design of AI Financial Research Workflows, arXiv:2608.24842v1, 25 August 2026.
- National Institute of Standards and Technology, The TEVV-Athlon Framework for Evaluating AI Systems, initial public draft announced 7 August 2026.
- National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1, 2024, official page updated 8 April 2026.
Test the system before you scale it
If your organisation has connected the documents but cannot yet show how material knowledge changes decisions, the next step is not another upload. It is a structured test of the decision logic.