every model spec’d & versioned · harness ✓ greenchangelog →

Writing · Toolkits

Vendor Risk Interview Kit for Model Vendors

Ten questions to ask any calculator or model vendor before signing. The good answers are the same across vendors; the ability to answer them is what varies.

By Worthune Staff · 2026-08-14

The interview is not about scoring the vendor. It is about surfacing which vendors have the artifacts a caller’s risk team will later ask for.

A caller evaluating a financial-math vendor for a durable dependency benefits from a structured interview. The ten questions below are the working set. They are designed to surface the specific artifacts a caller’s procurement, engineering, and risk teams will later ask for. Vendors that can answer all ten with pointers to public documents are durable partners; vendors that answer any of them with a promise instead of an artifact are worth watching before committing. The questions apply to Worthune and to any comparable vendor.

The ten questions

Q1. Where is the spec for each model published? A good answer is a URL that resolves to a human-readable specification with inputs, outputs, assumptions, and exclusions. A weak answer is a promise that the spec exists internally. Worthune’s answer: every model’s spec is public at worthune.com/docs/models/{model}, and every API response links its own spec and contract.

Q2. How does the vendor verify that its implementation matches the spec? A good answer is a described process — dual implementation, verification cases, release-gating on agreement. A weak answer is a claim of correctness without a process behind it. Worthune’s answer: each model is implemented twice — the production engine and an independent implementation built from the spec — with 250 verification cases per model, and releases gate on exact agreement; the eval datasets are public and free.

Q3. What is the source of the vendor’s regulatory constants, and how quickly does the vendor incorporate updates? A good answer names primary sources (IRS, SSA, state agencies) and describes an update process. A weak answer is a claim of currency without a described mechanism. Worthune’s answer: a public facts registry in which every constant carries its primary source (IRS revenue procedures and notices, SSA publications), its period, and a verified-on date — scoped to US-federal rules, and it says so rather than implying more.

Q4. What does a successful API response include? A good answer names the specific fields: model, spec version, inputs, outputs, sentinels, assumptions, facts, a cryptographic record, and a disclaimer. A weak answer is that the response contains the outputs and some metadata. Worthune’s answer: exactly that field list, in every successful response.

Q5. What is the vendor’s versioning policy? A good answer commits to versioning behavior changes and publishing a changelog. A weak answer treats versioning as an implementation detail the caller does not need to see. Worthune’s answer: behavior changes only with a spec version bump, recorded in the public changelog at worthune.com/models/changelog.

Q6. What is the vendor’s posture on backward compatibility and pinning? A good answer allows the caller to pin a specific spec version and guarantees the pinned version’s behavior. A weak answer moves silently under the caller. Worthune’s answer: version pinning is a paid-tier guarantee; on any tier, the versioned envelope plus the changelog make drift detectable the day it happens.

Q7. What happens if the caller stops using the vendor? A good answer includes portability commitments: specs are published, the caller’s stored envelopes remain interpretable, and the caller can migrate. A weak answer treats the caller’s data as vendor-owned. Worthune’s answer: the specs are public documents, and stored envelopes are self-describing — each carries its inputs, outputs, spec version, and the hash recipe to verify them — so they stay interpretable with or without the vendor.

Q8. What is the vendor’s pricing durability commitment? A good answer names, in writing, what stays free and what paid tiers exist to add. A weak answer describes current pricing without commitment to its shape. Worthune’s answer: a written commitment on the pricing page that the models, API, MCP server, and embeds stay free with attribution, with paid tiers existing to add guarantees, never to take back what is free.

Q9. What is the vendor’s recordkeeping story? A good answer includes the audit envelope, the SHA-256 record, and a described mechanism for later verification. A weak answer treats recordkeeping as the caller’s problem. Worthune’s answer: every response carries a SHA-256 record over the model, spec version, inputs, and outputs, with the verification recipe stated in the response itself.

Q10. What is the vendor’s incident-response process for correctness incidents? A good answer describes the workflow: disclosure, root-cause analysis, changelog entry, communication to affected callers. A weak answer has no described process because no incident has ever occurred (a claim that only means no incident has been noticed yet). Worthune’s answer: corrections ship as spec-version bumps with public changelog entries, and each spec documents its model’s known issues in the spec itself.

How to score the answers

The interview is not scored on a numeric scale. Each question is scored as artifact-backed, described-but-unverified, or promised. A vendor with ten artifact-backed answers is a durable partner. A vendor with a mix is worth watching to see whether the described answers become artifact-backed over time. A vendor with several promised answers is fine for short-term evaluation and risky for durable commitment.

The questions the interview does not include

The kit deliberately excludes questions about pricing dollar figures, about specific SLAs, and about specific customer references. Those are legitimate procurement questions, and they belong on a separate list run by the procurement team. The kit is scoped to the ten questions whose answers reveal how the vendor thinks about the discipline of shipping financial math; the procurement questions reveal how the vendor thinks about a commercial relationship, which is a different assessment.

Running the interview at scale

Callers evaluating multiple vendors benefit from running the same ten questions against each. The answers are directly comparable, and the artifact-versus-promise scoring produces a portfolio view that is more useful than any single vendor’s answers in isolation. The kit is designed to be repeated, and the value compounds as the caller’s comparison base grows.

The interview is not scored. It is documented. The documentation is what the risk team will need in twelve months.

What happens after the interview

A vendor whose answers scored well becomes a candidate for a durable commitment. A vendor whose answers scored partially is a candidate for a smaller commitment with a re-evaluation date. A vendor whose answers did not include the artifacts the interview asked for is a candidate for short-term evaluation with an explicit end date. In each case the caller has a specific plan; the interview’s value is not the conclusion but the specificity.

The interview and self-assessment

Vendors can run the interview against themselves. A vendor whose internal answers to the ten questions do not point to artifacts is a vendor whose durability commitment is a promise. The kit is not adversarial; it is a working list of the questions any caller will eventually ask, and vendors that maintain artifact-backed answers to all ten are the vendors that pass future audits with the least friction. Callers and vendors benefit from the same set of answers; the kit is scoped to the caller’s perspective, and vendors are welcome to use it as a self-check.

The follow-up cycle

The interview is not a single event. A vendor whose answers scored well at signing may drift; a vendor whose answers scored partially may improve. Re-running the ten questions annually keeps the caller’s picture of the vendor current. Callers who run the interview at signing and never again find that their picture of the vendor becomes historical rather than actual; the vendor may have changed, and the caller has no record of the change. Callers who re-run the interview annually catch drift in either direction and adjust the depth of their dependency accordingly.

The annual re-interview is short. The caller does not have to re-do the full evaluation; the caller has to confirm that the artifacts each answer pointed to are still in place, still current, and still resolve. A URL that resolved at signing and returns a 404 a year later is a signal. A described process that has changed materially is a signal. A pricing durability commitment that has been quietly modified is a signal. The signals are not always concerning, but they are worth knowing.

Sources

  1. [1] Worthune API documentation (published model specs). https://worthune.com/docs
  2. [2] Worthune pricing (free-tier commitment, in writing). https://worthune.com/pricing
  3. [3] Worthune public changelog. https://worthune.com/models/changelog