THE MISSING BLACK BOX FOR AI JUDGEMENT
AACIF
Every mature system that delegates consequential work builds in a declaration before reliance, and a record after it. Aviation does this. Banking does this. Medicine does this. AI is the exception.
The AI Assertion and Chain Integrity Framework requires consequential AI systems to declare what they knew, on what basis, and for whom — and requires that declaration to survive every handoff to the point where someone relies on the result. AACIF is a control framework, not a governance framework. It does not replace NIST, ISO 42001, or the EU AI Act. It produces the evidentiary substrate those frameworks assume already exists.
SEVEN QUESTIONS
The Seven Assertions
When an AI makes a call that affects someone's life, someone should be able to go back later and answer seven questions about it. Right now, for almost every AI system in use, nobody can answer any of them.
A1
Temporal
How old is what the AI knows?
An AI system learned what it knows by a certain date and stopped — but it answers every question in the same confident tone whether its knowledge is a week old or three years old. Example: DAX Copilot never tells a doctor how recent its medical knowledge is.
A2
Knowledge State
Which version of the rules was it using?
Laws, medical guidelines, and company policies all change. An AI system is built around a specific version of whatever rulebook it follows — and that version can go out of date without the AI knowing or saying so. Example: nH Predict's clinical guidelines update annually, but nothing in a denial letter says which year's version drove that specific decision.
A3
Population
Was this AI ever tested on someone like the person it's being used on?
A model can be accurate on average for the group it was trained on, and still be wrong most of the time for someone who doesn't match that group — older, sicker, or more unusual than the typical training case. The question is never “how accurate is this model,” it's “accurate for who.”
A4
Measurement
Is this actually confirmed, or has it just been repeated a lot?
An AI can sound certain about something because it saw the same claim many times in training data — not because that claim was ever independently checked. Example: Workday's applicant “fit score” sounds precise, but nothing requires the company to say whether it's based on checked outcomes or on patterns that simply repeat in old hiring records.
A5
Language / Perspective
Whose point of view is this, really?
If almost everything an AI learned from came from one language or one type of source, it can sound like broad agreement when it's really one narrow point of view, repeated often.
A6
Translation
What changed when the answer got turned into a decision?
An AI's raw output is often a probability — “73% confident,” for example. Somewhere along the way that soft number becomes something hard and final: approved or denied. That conversion is a real decision, made by someone, and right now nobody has to document that it happened or explain how. Example: MiDAS turned a simple mismatch in reported earnings — which could easily have been a clerical error — directly into a fraud accusation, with no declared step in between.
A7
Chain Integrity
Did all of the above survive long enough to matter?
This one is a check on the other six, not a new thing to declare. An AI system can get every one of A1 through A6 right at the moment it produces an answer, and still lose all of that information by the time a human being actually relies on the result — stripped out at every handoff, from the model, to the software system, to the case file, to the letter that finally reaches the person affected.
WHERE THE DECLARATION MUST TRAVEL
The Twelve-Layer Chain
AI is not one system to be assessed whole. It is a layered chain, and each layer has a different owner, a different failure mode, and a different recording obligation. The declaration must follow the output through all of it.
LAYER 0
Deployment Envelope
Tthe pre-authorized domain, authority level, and hard stops — set before training, governing every layer that follows
LAYER 1
Corpus
The raw and curated source material — the first place assertions are stripped
LAYER 2
Training
The process that turns corpus into behavior
LAYER 3
Model
The artifact at a point in time — identifier, version, release date
LAYER 4
Instruction
System prompts, guardrails, tool permissions
LAYER 5
Retrieval
Live knowledge at inference — documents retrieved, source currency
LAYER 6
Inference
Computation at the moment of use — currently the least governable layer in the chain
LAYER 7
Output
The recommendation, score, or instruction, with all assertion declarations attached
LAYER 8
Integration
The handoff into the enterprise system — whether the declaration survived the transformation
LAYER 9
Human Review
Where a person accepts, rejects, or ignores the output
LAYER 10
Action
The real-world decision or execution
LAYER 11
Audit / Recovery
Later reconstruction, claims, litigation — whether the chain can be replayed from action back to source
A system can pass every component-level review and still fail because the handoffs stripped the information needed to understand the judgment. AACIF treats that loss as a breach — not a documentation inconvenience, but a failure of the layer that was supposed to carry the evidence.
LIVE TOOLS
See the framework, not just read it
Two working instruments make AACIF's structure interactive rather than theoretical — the chain a declaration has to survive, and the taxonomy that governs how AI outputs get classified along the way.
⬤ AACIF INSTRUMENT
AI Chain Visualizer
The Constitution is the mission statement. The Preamble is the stated purpose. Five years of research — organized across six analytical pillars — are the work papers. It is the shared foundation beneath three published works below. If you want to audit the auditor, this is the starting point.
⬤ AACIF INSTRUMENT
AACIF Taxonomy Architecture
The classification structure behind the seven assertions — how AI outputs get sorted (fact, inference, prediction, judgment) and how that classification is supposed to travel with the output.
FIELD VALIDATION
Ten systems settled. An eleventh in progress.
A governance framework that cannot be applied to real systems is a working hypothesis. AACIF has been applied to ten production AI systems currently operating across seven sectors — healthcare, employment, financial services, legal, professional services, supply chain, and government benefits. All ten were selected for one shared trait: consequential AI output with no genuine human review before action. All ten are Tier 3.
​
The method was the same for each: assess against all seven assertions using only public documentation, and ask three questions. Is the required declaration present? Does it survive the handoffs to the point of action? Can an independent expert later reconstruct what the system knew at the moment it acted?
The answer, across all ten systems, at the chain level, was no — every time.
01
nH Predict (Optum / UnitedHealth)
90% of denials reversed on appeal; only 0.2% of patients ever appeal
02
DAX Copilot (Microsoft / Nuance)
audit trail deleted by design at 30 days, across 400+ health systems
03
Epic ART / Emmie
AI answers patient messages autonomously — the patient is never told AI answered
04
Workday Illuminate / HiredScore
documented 2022 failure operating outside validated parameters, no disclosure required
05
FICO Falcon Fraud Manager
documented 2022 failure operating outside validated parameters, no disclosure required
06
Upstart AI Lending
documented 2022 failure operating outside validated parameters, no disclosure required
07
Harvey AI
self-reported 0.2% hallucination rate vs. an independent finding of 1 in 6 queries
08
Harvey at PwC / Leah (ContractPodAi)
a 10-step agentic workflow where an error at step 2 propagates silently to the final output
09
Blue Yonder Demand AI
a 7+ system pipeline where the confidence interval is stripped at the very first handoff
10
MiDAS (Michigan)
40,000+ citizens falsely accused of unemployment fraud, 2013–2015, 85% error rate, $20M settlement in 2024
11
MARVEL (I-24, Nashville)
TIER PENDING
the first genuine reinforcement-learning system in the corpus, controlling posted speed limits across 67 gantries over 17 miles of I-24, ~160,000 daily commuters. Trained once in simulation and deployed as a frozen policy since March 2024. Its safety-guard override rate (~98% autonomous / ~2% guard-overridden) is the most transparently documented human/rule-based check of any system assessed. Unlike the ten systems above, MARVEL's tier is withheld rather than settled: whether it belongs in Tier 2 (genuine human override) or Tier 3 (override present but practically defeated) depends on one unresolved fact — whether the Traffic Management Center's override authority is real or only vendor-asserted. Not yet independently confirmed with TDOT or Vanderbilt.
The most important finding is not that these systems fail their governance requirements. It's that they satisfy them. NIST, the EU AI Act, ISO 42001, and Singapore's MGF all require documentation. None require the specific declarations AACIF defines. None require those declarations to survive a handoff. None make the chain itself the unit of governance. The existing frameworks did exactly what they were built to do — the problem is what they were built to do.
​
One more system is in active assessment, not yet complete: a new line of work assessing autonomous vehicles against AACIF — including real litigation outcomes — is underway and will be added to this page once complete. (MARVEL, previously listed here as a pending item, is now entry 11 in the ledger above, with its tier explicitly withheld rather than assigned.)
WORKING GROUP
Built by people who've done this before
Thom Barrett
Joined Coopers & Lybrand in 1978 in a new group focused on EDP auditing. Became the subject matter expert on operational controls across transfer agency, fund administration, and compliance. Introduced XBRL to the SEC. Retired from the PwC partnership in 2012.
Michael Willis
More than thirty years as a PwC partner. Served alongside Barrett on PwC's Innovation Working Group and as technical lead on the XBRL team. Founding Chair, XBRL International. Former Associate Director, SEC Division of Economic and Risk Analysis.
J. Donald Warren Jr.
Retired PwC partner, worked alongside Barrett for more than twenty years. Built continuous auditing as a discipline — former Director, Center for Continuous Auditing, Rutgers; current Chair, Lamar University. Domain: the control-layer architecture.
John W. Stadtler
Active PwC partner, Governance Insights Center. Domain: corporate governance and board oversight — how AACIF connects to COSO's generative-AI guidance, giving boards evidence rather than vendor-produced documentation.
Matthew Chafee, PhD
Professor of Neuroscience, University of Minnesota — cognitive control, working memory, prefrontal cortex. Domain: the human layer, and what genuine independent review requires neurologically.
Scott E. Eston
Retired COO and Executive Committee Chairman, GMO; independent trustee, Eaton Vance Funds. Began at Coopers & Lybrand alongside Barrett as assurance partner to his technology partner. Domain: the manager due-diligence analogy at the framework's core.
HONEST SELF-CRITIQUE
Stress-tested, not defended
Before AACIF goes further, it is being stress-tested against real domain expertise — not defended, mapped. A set of research notes name specifically where the framework is most vulnerable, prepared as briefs for conversations with a computer scientist, a neuroscientist, a regulator, and others whose expertise the framework depends on but does not yet have.
​
These notes ask, among other things: is boundary-contract governance at the AI inference layer technically achievable at all, given that inference remains the least inspectable layer in the entire stack? What happens to human review when an AI is right often enough that trusting it becomes the rational choice — including in exactly the cases where it's most likely to be wrong? Who is missing from the table where this standard gets set, and why must the capital that will be governed by the standard be excluded from writing it?
​
AACIF does not yet have all the answers. It is precise about which questions remain open, and to whom they are addressed.
What AACIF is producing
AI Risk and Controls: The Assembly Line Nobody Inspected
Alongside the framework itself, this research is producing a second book — applying the same evidentiary discipline directly to the risk-and-controls profession. It names the standards gap in NIST's voluntary framework, the EU AI Act's documentation-without-verification model, and ISO 42001's process-without-substance certification, and proposes six specific, auditable controls and a tiered risk classification for consequential AI deployment, addressed to the people who already have the authority to require it.
