every model spec’d & versioned · harness ✓ greenchangelog →

Writing · MCP & AI assistants

The Evidence Grid: Tool-Backed Numbers vs. Remembered Numbers

Every claim in this piece is checkable: real envelope output on one side, the registry's own historical rows on the other. What two cycles of constant drift actually do to an assistant that answers from memory.

By Worthune Staff · 2026-08-14

The case for tool-backed numeric answers is usually argued in principle. This piece argues it with receipts — a real computed envelope, and the registry's historical rows standing in for what stale memory looks like.

The argument that assistants should compute rather than recall is easy to nod along with and easy to discount, because it is usually made with hypotheticals. So this piece does something stricter. Every number in it comes from a checkable artifact: a live model run whose envelope is quoted with its hash, and the facts registry's own rows — current and historical — which show precisely how far constant drift moves the numbers a remembered answer depends on. No invented assistant transcripts, no staged failures; just the two data sources, side by side.

Exhibit one: what a tool-backed answer carries

Run the employer-match model with a $120,000 salary, a 6 percent employee contribution, and a 50-percent-up-to-6-percent match, and the envelope returns: employee contribution $7,200, employer match $3,600, total annual savings $10,800, computed under spec version 1.0.0. The interesting part is not the arithmetic — it is what rides along. The facts array cites irs.401k.elective-deferral.2026 at $24,500 and irs.401a17.comp-limit.2026 at $360,000, both sourced to IRS Notice 2025-67, because those limits bound the computation. The assumptions array states the match formula and the caps in prose. And the record carries a SHA-256 — this run's begins 1ec1d6a9 — over the model, spec version, inputs, and outputs. An assistant relaying this answer can cite its version, its constants, and their sources, because the tool result contains all three.

Exhibit two: what drift does to remembered constants

A model trained before the latest fall announcements carries values at least one cycle stale — and the registry's TY2024 rows record what a stale vintage actually held. The grid below is real data on both sides: the 2024 rows the registry keeps for the record, against the 2026 rows the models compute with today.

ConstantTY2024 (historical rows)TY2026 (verified rows)
415(c) overall limit$69,000 — IRS Notice 2023-75$72,000 — IRS Notice 2025-67
Estate basic exclusion$13,610,000 — Rev. Proc. 2023-34$15,000,000 — Rev. Proc. 2025-32 / OBBBA
Top of 10% bracket, single$11,600 — Rev. Proc. 2023-34$12,400 — Rev. Proc. 2025-32
Top of 10% bracket, joint$23,200 — Rev. Proc. 2023-34$24,800 — Rev. Proc. 2025-32

Read the grid as an error budget. An assistant answering a 2026 mega-backdoor question from 2024 memory is $3,000 off on the governing limit before it says a word about strategy. An estate-planning answer from the same vintage misses the exclusion by $1.39 million. None of these errors announces itself — a remembered $69,000 sounds exactly as confident as a computed $72,000, and the user cannot tell the difference from the prose. The registry can, because dated, sourced rows are the one place where "which year is this number from" has a mechanical answer.

The grid as a standing method

The two exhibits compose into a test any team can run quarterly, with no staging. Take the constants your product's answers depend on; pull their current and prior rows from the registry; ask your assistant each value without tools and grade against the verified rows; then ask again with tools and confirm the envelope's facts array cites the current rows. The first half measures your exposure to drift; the second half verifies the plumbing that eliminates it. Both halves grade against public, dated artifacts — which means two teams running the same grid get comparable results, and a compliance reviewer can re-run either half without trusting anyone's screenshots.

What the evidence does and does not claim

The grid shows that constants drift materially year over year, that remembered answers cannot signal their own vintage, and that tool-backed answers carry their vintage — version, constants, sources — in the response itself. It does not claim any particular assistant fails any particular prompt; that is what the red-team prompt set (/writing/red-team-prompts) helps you measure for your own stack. The division of labor is deliberate: this piece supplies the checkable baseline, your evaluation supplies the verdict about your system, and neither has to borrow credibility from the other.

Sources

  1. [1] Worthune facts registry (current and historical rows). https://worthune.com/facts
  2. [2] IRS Notice 2025-67 (TY2026 retirement plan limits). https://www.irs.gov/pub/irs-drop/n-25-67.pdf
  3. [3] Worthune writing: Red-Team Prompts for Financial Assistants. https://worthune.com/writing/red-team-prompts