When AI Refuses to Guess: A Month-End Evidence Packet

Treasury Desk Brief · Before the Journal Entry
Public-source and synthetic method note Primary scope: month-end / reconciliation evidence Source pack retrieved 2026-08-02 Human-reviewed; no posting or close approval

A reviewer-first evidence packet for source lineage, missing data, rule and model versions, exceptions, human decisions and month-end sign-off.

An AI-assisted month-end workflow can make the responsible choice and still leave the reviewer with something useless: a blank no one can explain.

A blank might mean a required source was missing. It might also mean a parser failed, a rule did not cover the case, the request exceeded authority, or the system returned nothing. An estimate can look cleaner while hiding the same evidence gap. A generic error is visible, but it rarely tells a controller which records were affected, what the system tried to do, or what should happen next.

Core distinction. A refusal is reviewable only when the record explains the non-action, the evidence gap, the owner, and the next permitted step.

A refusal can still leave an unusable blank

A reviewable refusal is different. It records a deliberate non-action or non-inference under a defined policy. The record connects the task to its source, scope, authority, reason, owner, and next permitted step.

Month-end work is a chain, not a single action: task authorization → source intake → completeness check → rule or model execution → refusal or candidate output → exception queue → human decision → downstream handoff → sign-off. The record needs to preserve those state changes, not just the final spreadsheet.

The NIST AI Risk Management Framework 1.0 supports documenting knowledge limits, defined scope, output context, and human-oversight processes. The NIST Generative AI Profile adds provenance, model and version inventories, retention, change history, and differentiated human roles. Neither source prescribes a finance-specific packet. Together, they support the discipline behind one: make the boundary visible before someone relies on the output.

Before the run: define the evidence contract

A refusal only means something if the workflow first defines what counts as a required input and what the system is authorized to do.

Before execution, the run record should identify the close period, entity or business-unit scope, task scope, source systems, expected population, relevant control totals, and permitted action class. “Propose matches” is not “post entries.” “Summarize exceptions” is not “resolve exceptions.” The authority record should make those differences explicit.

The logic needs the same treatment. Rules, tolerances, mappings, and ownership need controlled versions before automated application can be reviewed consistently. A model version may matter. A prompt or instruction version matters only where it can materially change the result. A file hash or immutable reference may identify a snapshot, but it does not establish that the underlying business evidence is correct.

COSO’s Achieving Effective Internal Control Over Generative AI offers relevant design examples for automated transaction processing and reconciliation: contextual exception queues, configuration controls, override rationales attached to transaction records, separation-of-duties considerations, and monitoring of errors, reversals, and exception clearance. Those examples can shape a review record. They do not show that any particular implementation operates effectively.

Required, conditional, and proposed fields

Not every field belongs in every run. The packet should say which fields are required, conditional, or proposed instead of presenting a long schema as a universal requirement.

Required fields define the minimum evidence contract for the task: run identity, scope, source references, authority, action state, and any human decision or sign-off gate that applies. When a refusal occurs, the code, reason, affected records, and exception owner are required because they explain the non-action.

Conditional fields depend on materiality and workflow design. A prompt or instruction version is useful where it can change the result; it is noise in a purely deterministic step. A model version matters only when a model participates. Control totals matter when the process has a defined expected population. An override record matters when someone departs from a candidate or control path.

Proposed fields may improve integrity or operating clarity but are not imposed by the cited sources. A hash might identify exact bytes; a controlled immutable reference may be more appropriate in a managed system. Either can support lineage. Neither proves the source was authentic, complete, or correctly interpreted.

The better review question is not “Does the packet contain every possible field?” It is “Does the packet contain enough evidence to understand this authorized task and its unresolved states?”

At the refusal: record what did not happen

A useful refusal record answers six questions:

  1. What was requested? The exact task or action, not a generic process label.
  2. What was not done? No estimate, match, adjustment, posting, or another defined non-action.
  3. Why? Missing required source, stale source, conflicting evidence, incomplete scope, unauthorized action, missing rule, schema mismatch, unsupported inference, low-confidence candidate, or system failure.
  4. What evidence is missing or conflicting? The exact source object, field, batch, period, entity, or rule.
  5. Which records are affected? The population and downstream dependency should remain visible.
  6. Who owns the exception? A role, due date, escalation path, and current state turn the refusal into managed work.

These terms are not interchangeable. A refusal is a deliberate non-action under a defined condition. An abstention is a defined non-answer or non-decision. An error is a tool or system failure. An exception is the managed work item. An escalation routes the item to someone with more authority or expertise. A blocked action enforces an authority or policy boundary. An unresolved candidate may exist before an authorized disposition.

A silent blank preserves none of this. A generic error preserves almost none of it. A reviewable refusal preserves both the reason and the route forward.

After the refusal: separate the candidate from the human decision

A refusal does not have to stop all useful work. The workflow may process complete records, prepare a limited candidate, identify alternatives, or assemble the evidence needed for a human decision. The packet should show which output is permitted and which action remains blocked.

The human record is separate from the AI-generated candidate. “Human approved” is too thin. The record should identify the reviewer role, evidence inspected, decision, rationale, any override, unresolved items, and timestamp. Human gates matter most for judgment, overrides, unresolved exceptions, and sign-off. They are not a universal requirement for manual approval of every deterministic step.

An override should preserve the earlier candidate or rule result, the authority to depart from it, the reason, and the downstream effect. A reviewer decision should say accept, reject, modify, or defer. Sign-off should identify the scope approved and the items still open. These records make judgment visible; they do not make it correct.

Downstream: show what left the workflow

Month-end evidence often breaks at the handoff. The run log may be complete while the downstream workbook or close queue contains a different value. The packet should identify the output file or controlled reference, destination, receiving owner, generated time, acceptance state, and restrictions.

Change history matters too. If a source is re-exported, a rule is corrected, or a human changes a candidate, the new run should link to the superseded run and preserve the reason and approver. Reproducibility means a reviewer can reconstruct or rerun the identified inputs and configuration. It does not mean the inputs were complete, the rules were appropriate, or the result was accurate.

PCAOB AS 1215: Audit Documentation is used here only as a documentation analogy within PCAOB audit engagements. Its concepts—linking purpose, source, work performed, evidence, conclusions, performer, reviewer, and dates—show why a later reviewer needs more than a polished output. AS 1215 does not make this packet a universal finance-AI requirement and does not provide audit assurance for the workflow described here.

Month-End Evidence Packet

Synthetic illustrative design — not a real company record and not a universal audit requirement.

Packet sectionCore evidence fieldsReviewer purposeWhat it does not proveStatus
Run_MetadataRun ID; close period; entity/task scope; start/generated time; timezoneEstablish the defined run and link rerunsCorrectness or accounting completenessRequired
Source_RegisterSource system/owner; record ID; snapshot/version; retrieved-at; freshness ruleTrace point-in-time inputsSource accuracy or completenessRequired
Input_SnapshotExpected/received counts; control total; filters; exclusions and reasonsExpose population and completeness checksFinancial-statement assertionsConditional by workflow
Authority_RecordRequester role; authorized action; policy version; segregation checkShow what the system may doLegal authority outside organizational policyRequired
Rule_Model_RegisterRulebook/mapping/tolerance versions; model/access mode; prompt version where materialReconstruct the logic and material configurationRule appropriateness or model reliabilityRequired / conditional
Action_LogRequested, attempted, and performed action; source fields relied on; records affectedSeparate intent, attempt, and outcomeAn accurate resultRequired
Refusal_Exception_LogRefusal code/reason; missing or conflicting evidence; owner; due date; escalation; statusExplain non-action and create accountable workGood control design or resolutionRequired when refusal occurs
Candidate_OutputsCandidate; alternatives; uncertainty note; dependencies; excluded recordsSupport a defined reviewCorrectness or approvalConditional
Reviewer_DecisionsReviewer role; evidence inspected; decision; rationale; override; unresolved itemsRecord the human dispositionProfessional correctness or independenceRequired at a human gate
Downstream_HandoffOutput ID; destination; receiving owner; accepted/rejected state; restrictionsTrace what left the workflowPosting, reconciliation, or close completionRequired if exported
Human_SignoffSign-off scope/role/time; open items; explicit non-sign-offs; reopen conditionsBound the approved stageAudit opinion or completed closeRequired at a defined gate
Change_Log / Integrity_RecordChanged field; prior/new reference; reason; approver; superseded run; file ID/hash or controlled reference where appropriatePreserve lineage and identify the artifactAuthenticity or correctness of business evidenceRequired for material changes; integrity field proposed / conditional

Organizations will not need every field in the same form. The point is to make the evidence contract explicit, so required inputs and decisions do not disappear inside a model output.

Synthetic walkthrough: one missing settlement batch

Synthetic illustrative data — not a real company record.

A finance team runs a month-end match between a payment-processor settlement export and an ERP cash-clearing ledger. The workflow may propose matches and exceptions. It may not post entries or approve the close.

Evidence fieldSynthetic valueReview meaning
Task scopePropose settlement-to-ledger matches; no postingDefines the authority boundary
Expected / received batches240 expected; 239 receivedMakes the completeness exception visible
Missing sourceProcessor batch P-0731-188Identifies the exact required evidence object
Rulebook versionrecon_rules_2.7Shows which deterministic rules were authorized
Model useModel summarizes exceptions onlySeparates summarization from matching authority
Refusal codeR-MISSING-REQUIRED-SOURCERecords a defined non-inference state
Action not takenNo estimate; no proposed clearing adjustmentShows what the system deliberately avoided
Permitted partial outputProcess 239 complete batches; carry one exceptionPreserves useful partial work
Exception ownerPayments Operations ManagerCreates accountability without personal data
Human decisionRequest source re-export; accept partial packet; reject estimated adjustmentRecords the substantive disposition
Downstream handoffReview file sent to the Close Exception QueueLinks the output to the next owner
Sign-off boundaryMatching work reviewed; reconciliation and close sign-off remain openPrevents status inflation

The example shows a specific discipline: the system refuses one unsupported inference, preserves the missing evidence, and routes the exception without pretending that the reconciliation or close is complete.

Evidence is not assurance

Evidence object or stateNon-equivalence
Log exists≠ log is complete
Explanation exists≠ decision is correct
Human approval exists≠ substantive review occurred
Reproducible run≠ accurate result
Control design≠ operating effectiveness
Audit trail≠ audit assurance
AI refusal≠ system safety
Exception visibility≠ exception resolution
Reviewer-ready packet≠ completed close
Matching hash≠ correct business evidence

The packet is designed to support reviewer reconstruction. It is not evidence of compliance status, control operating effectiveness, production deployment, auditor acceptance, or capability for autonomous reconciliation, posting, or close approval.

What the human reviewer still decides

The packet does not decide accounting treatment. It does not establish financial-statement assertions. It does not determine whether a control operated effectively over a period or population. It does not replace the controller, accountant, internal auditor, external auditor, legal, tax, or compliance owner.

The human reviewer still decides whether the source is sufficient, the scope is complete, the rule is appropriate, the candidate is supportable, an override is justified, the exception can be closed, and the downstream stage can be signed off. The packet makes those decisions visible. It does not make them correct.

Reader-facing sources

Source retrieval dates are preserved from the evidence pack. Exact links were checked again during clean-copy assembly.

  1. Artificial Intelligence Risk Management Framework (AI RMF 1.0) — NIST. Published 26 January 2023; retrieved 2 August 2026.

    Supports: Knowledge limits, bounded scope, output context, documentation, and human-oversight processes. Does not prove: A finance-specific packet, accounting treatment, control operating effectiveness, correctness, or audit assurance.

  2. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile — NIST. Published 26 July 2024; source page updated 8 April 2026; retrieved 2 August 2026.

    Supports: Provenance, model/version inventories, retention, change history, differentiated human roles, and override/incident records. Does not prove: A mandatory month-end schema or proof of output correctness.

  3. Achieving Effective Internal Control Over Generative AI — COSO. Released 23 February 2026; retrieved 2 August 2026.

    Supports: Relevant control-design examples for transaction processing, reconciliation, exception queues, overrides, configuration, and monitoring. Does not prove: Implementation, operating effectiveness, compliance, auditor acceptance, or assurance.

  4. AS 1215: Audit Documentation — PCAOB. Current text retrieved 2 August 2026; source page notes amendments effective 15 December 2026; retrieved 2 August 2026.

    Supports: Reviewer-reconstructability concepts inside PCAOB audit engagements: purpose, source, evidence, conclusions, performer/reviewer, and dates. Does not prove: A universal finance-AI requirement, a universal retention rule for this packet, or audit assurance.

  5. SP 800-53 Rev. 5: Security and Privacy Controls for Information Systems and Organizations — NIST. Rev. 5 / current update page; retrieved 2 August 2026.

    Supports: A bounded analogy for event-record fields such as event type, time, source, outcome, and identity. Does not prove: Accounting correctness, month-end completion, or audit assurance.

Public caveat. This is a public-source and synthetic method note. The Month-End Evidence Packet is an illustrative reviewer-oriented design, not a universal audit checklist. It is not accounting, audit, legal, tax, compliance, investment, or cybersecurity advice or assurance. It does not establish control operating effectiveness, completed reconciliation, journal posting, close approval, or a verified Treasury Desk production workflow.

Which field would your team require in an AI-assisted month-end run log? Please share only public methodology, synthetic examples, or public-source corrections—not private close records or control deficiencies.

Leave a Comment